GraphQL is often described as a query language, which undersells what you are actually building. A GraphQL service is a typed contract (the schema) plus an execution engine that walks a client's query field by field and calls a function, a resolver, for each one. The client decides the shape of the response, and the server decides how much any shape is allowed to cost. Almost every operational problem a GraphQL deployment has comes from forgetting the second half of that sentence.

This page works from the request inward: execution, batching, cost limits, caching and federation. For the REST and RPC alternatives see REST API design and gRPC and Protobuf.

The request path in one picture

Clientsweb, iOS, AndroidCDN / edgeGET by query hashGateway or routerparse, validate, cost check, planquerymissExecution engineresolver per field, null propagationOperation registrytrusted documents, cost limitsallowlistProducts subgraph@key(id)Reviews subgraphextends ProductPer-request DataLoadersbatch and dedupe by keyDatabases and servicesone batched call per levelload(keys)The client chooses the shape; the gateway decides what any shape may cost.
A query passes through the edge and gateway, where it is validated and costed, before resolvers execute; per-request loaders turn per-object lookups into one call per level.

Parse, validate, execute

Every request goes through three phases. Parse turns the query text into an abstract syntax tree. Validate checks that tree against the schema: fields exist, arguments and variables have the right types. A query that fails validation never executes. Execute walks the selection set from the root operation type (Query, Mutation or Subscription) and calls the resolver for each field, passing the parent object, the field arguments, a per-request context and some metadata about the path.

A resolver is just a function that returns a value or a promise. If the field is a scalar the value goes into the response; if it is an object type, execution recurses into its sub-selection with that object as the new parent; if it is a list, it recurses once per element. Fields on a query run in any order and may run concurrently; top-level fields on a mutation run strictly in sequence, so a client can rely on the second mutation seeing the first one's effects.

Errors do not abort the response. A throwing resolver's field becomes null and gets an entry in the errors array with a message and a path such as ["product", "reviews", 3, "author"]. If the field was declared non-null (User!) it cannot be null, so the null propagates to the nearest nullable parent, possibly wiping out a large part of the response. That is the most important schema-design consequence of the execution model: mark a field non-null only when the server can genuinely always produce it, and keep fields backed by other services nullable so one failing dependency degrades a branch rather than the whole page.

Designing a schema you can evolve

The schema is the public contract, and it is harder to change than a REST resource because you cannot see which clients select which fields without instrumentation. A few conventions make it evolvable. Use cursor-based connections for every list that can grow, so pagination is uniform and the server controls page size. Give each mutation a single input type and a result type, so you can add fields to either without breaking callers. Model expected business failures as data, with a union of success and error types, and reserve the errors array for unexpected faults; a client then handles ValidationFailed with an exhaustive switch instead of parsing message strings. Retire fields with @deprecated and delete them only when usage metrics reach zero.

type Query {
  product(id: ID!): Product
  products(first: Int = 20, after: String): ProductConnection!
}

type Product {
  id: ID!
  name: String!
  price: Money!
  reviews(first: Int = 5, after: String): ReviewConnection!   # backed by another service
}

type ProductConnection { edges: [ProductEdge!]!  pageInfo: PageInfo! }
type ProductEdge { cursor: String!  node: Product! }
type PageInfo { hasNextPage: Boolean!  endCursor: String }

type Mutation { addReview(input: AddReviewInput!): AddReviewResult! }

input AddReviewInput { productId: ID!  rating: Int!  body: String! }
union AddReviewResult = AddReviewSuccess | ValidationFailed | ProductNotFound
type AddReviewSuccess { review: Review! }
type ValidationFailed { field: String!  message: String! }

The N+1 problem and batching

Because each field resolves independently, the obvious implementation issues one backend call per object. Ask for 20 products with their reviews and each review's author and the engine calls the reviews resolver 20 times and the author resolver once per review. That is the N+1 problem, and in GraphQL it compounds with every level of nesting.

The fix is a batching loader. Resolvers call loader.load(key); the loader collects every key requested during one tick of the event loop, calls a batch function once, and memoises results within the request. The batch function must return results in the same order and length as the keys, with an error in place of any missing item. Create loaders per request: a process-wide loader leaks one user's data to another.

import DataLoader from "dataloader";

// One set of loaders per request: the cache must never leak between users.
export function makeLoaders(db) {
  return {
    reviewsByProduct: new DataLoader(async (productIds) => {
      const rows = await db.query(
        "SELECT * FROM reviews WHERE product_id = ANY($1) ORDER BY created_at DESC",
        [productIds]);
      const byId = new Map(productIds.map((id) => [id, []]));
      for (const r of rows) byId.get(r.product_id).push(r);
      return productIds.map((id) => byId.get(id));     // same length, same order as keys
    }),
    userById: new DataLoader(async (ids) => {
      const rows = await db.query("SELECT * FROM users WHERE id = ANY($1)", [ids]);
      const byId = new Map(rows.map((u) => [u.id, u]));
      return ids.map((id) => byId.get(id) ?? new Error(`user ${id} not found`));
    }),
  };
}

const resolvers = {
  Product: { reviews: (product, args, ctx) => ctx.loaders.reviewsByProduct.load(product.id) },
  Review:  { author:  (review, args, ctx) => ctx.loaders.userById.load(review.authorId) },
};

Batching changes the cost model from one call per object to one call per level of the query. Other ecosystems do the same under different names; Caliban's ZQuery is the Scala equivalent.

Worked example: one product page, three ways

Take a product listing page that sends this query: products(first: 20) { edges { node { name reviews(first: 5) { edges { node { rating author { name } } } } } } }. Suppose every product has at least five reviews, and the 100 reviews shown were written by 63 distinct users.

ImplementationBackend callsWhy
Naive resolvers1 + 20 + 100 = 121one list call, one reviews call per product, one user call per review
Per-request DataLoader1 + 1 + 1 = 3one call per level; the user batch carries 63 distinct ids, not 100

Both versions have three dependent levels, so on an idle system latency looks similar. Under load it does not: with 5 ms per query and a pool of 10 database connections, the naive version queues its 20 review queries in two waves and its 100 author queries in ten, about 65 ms in total, against about 15 ms for the loader version, and it sends the database forty times as many queries. The same query also gives a useful cost estimate before execution: 20 products, plus 20 times 5 reviews, plus 100 authors is 220 objects. That number is what a cost limiter should reason about, and it is computable from the query and its arguments alone.

Bounding what a query may cost

A public GraphQL endpoint accepts arbitrary programs, so bound them before execution. Depth limits reject deeply nested queries cheaply but crudely. Static cost analysis multiplies list page sizes down the tree, as in the example, and rejects anything over a budget; a rate limiter refilled in cost units then treats one 5,000-object query and fifty 100-object queries alike. Mandatory pagination makes the cost computable at all. Execution timeouts catch what static analysis misjudges.

def query_cost(selection, multiplier=1, defaults=None):
    """Static cost: objects the query can return, computed before any resolver runs."""
    total = 0
    for field in selection.fields:
        if field.is_list:
            n = field.args.get("first") or field.args.get("last")
            if n is None:
                raise QueryRejected(f"{field.name} must be paginated with first/last")
            if n > 100:
                raise QueryRejected(f"{field.name}: page size {n} exceeds 100")
            count = multiplier * n
        else:
            count = multiplier
        total += count * field.weight            # weight > 1 for expensive resolvers
        if field.selection:
            total += query_cost(field.selection, count)
    return total

cost = query_cost(parsed.operation.selection)
if cost > budget_for(client):                    # e.g. 5,000 for a public app
    raise QueryRejected(f"query cost {cost} exceeds budget")
ctx.rate_limiter.charge(client, cost)            # bucket refilled in cost units, not requests

For first-party clients, go further with trusted documents: at build time, extract every operation the app sends, register it with the server under its hash, and reject anything not on the list in production. Arbitrary queries then become impossible for attackers and every operation in production has been reviewed. Field-level authorisation belongs in resolvers or a directive layer, not only at the endpoint, since one query can touch many objects with different owners; broken object-level authorisation is as common in GraphQL as in REST, as the OWASP API Top 10 notes.

Caching when everything is one endpoint

HTTP caching assumes a URL identifies a resource. A GraphQL endpoint is one URL that receives POST bodies, so CDNs and browsers cache nothing by default. Three layers recover most of what REST gets for free. First, persisted queries replace the query text with its SHA-256 hash. Apollo's automatic persisted queries send extensions.persistedQuery = {version: 1, sha256Hash}; on a miss the server answers PERSISTED_QUERY_NOT_FOUND and the client retries with the full text. Hashed queries fit in a GET URL, so a CDN can cache public responses keyed by hash and variables, with a Cache-Control header computed from the shortest-lived field in the response.

# First attempt: hash only, as a cacheable GET
GET /graphql?extensions={"persistedQuery":{"version":1,"sha256Hash":"ecf4edb4..."}}&variables={"id":"p42"}
-> {"errors":[{"message":"PersistedQueryNotFound",
                   "extensions":{"code":"PERSISTED_QUERY_NOT_FOUND"}}]}

# Retry once with the full text; the server stores hash -> query
POST /graphql  {"query":"query Product($id: ID!) { ... }", "variables":{"id":"p42"},
                "extensions":{"persistedQuery":{"version":1,"sha256Hash":"ecf4edb4..."}}}

# Every later request from any client is a short GET the CDN can cache

Second, normalised client caches store objects by type name and id, so a mutation returning the updated Product refreshes every view of it. Third, server-side caching belongs at the data-source level below the resolvers; whole-response caching rarely pays because clients seldom send identical selections.

Federation: one graph, many teams

Federation lets each team own a subgraph while a router composes them into one supergraph. An entity such as Product declares @key(fields: "id"); any subgraph that knows the key can add fields to it. The router plans each query: fetch products, send their key representations to the reviews subgraph's _entities field in one batch, then merge. It is a DataLoader across services, and the plan's depth is a latency floor: three dependent hops cost three round trips.

# products subgraph
extend schema @link(url: "https://specs.apollo.dev/federation/v2.3", import: ["@key"])
type Product @key(fields: "id") { id: ID!  name: String!  price: Money! }

# reviews subgraph: contributes a field to a type it does not own
extend schema @link(url: "https://specs.apollo.dev/federation/v2.3", import: ["@key"])
type Product @key(fields: "id") { id: ID!  reviews(first: Int = 5): ReviewConnection! }

# What the router sends to reviews after fetching products:
query ($r: [_Any!]!) { _entities(representations: $r) { ... on Product { reviews { ... } } } }
# variables: {"r": [{"__typename": "Product", "id": "p1"}, {"__typename": "Product", "id": "p2"}]}

Federation adds a composition check in CI, a schema registry and a router that is now critical infrastructure. It pays off when several teams ship on one graph; with one team, a single modular schema is simpler.

The HTTP transport and status codes

GraphQL does not define a transport. The GraphQL-over-HTTP specification, still a draft but already implemented by several major servers, settles the common case. Clients should accept application/graphql-response+json. With that media type, 200 means the document passed validation and execution was attempted, even if the response contains field errors, and 400 means the request was malformed or failed validation and nothing ran. Monitoring must therefore read the errors array, not just status codes, or a resolver that fails on every request looks perfectly healthy. Subscriptions usually run over WebSockets or server-sent events. Incremental delivery with @defer and @stream exists in some servers but is still a proposal rather than part of the released specification, so check your server and client libraries before designing around it.

Failure modes

  • N+1 under a new field. Someone adds a resolver that queries directly; latency is fine in development with three rows and collapses in production. Alert on backend calls per operation, not just latency.
  • Non-null cascades. A non-null field backed by a flaky service nulls an entire list when it fails. Make cross-service fields nullable.
  • Invisible errors. Status 200 with a populated errors array hides failure. Emit per-field error metrics.
  • Breaking removals. A field deleted while an old mobile build still selects it breaks that build forever. Track field usage by client version before removal.

Trade-offs

GraphQL wins when many clients with different screens read overlapping, related data: it removes over-fetching and round trips and gives a typed contract. It costs you HTTP caching, simple rate limiting and load analysis from URL logs. REST suits public resource APIs and file transfer; gRPC suits internal calls with tight latency budgets. Many teams use GraphQL as an aggregation layer for their own apps, essentially a typed backend for frontend, over REST or gRPC services, which plays to its strengths and keeps it off the hot internal path.

What to do next

  1. Instrument backend calls per operation and find your worst N+1 offender.
  2. Put every read behind a per-request batching loader that preserves key order.
  3. Require first/last on every list field, cap page sizes, and add static cost analysis with a per-client budget.
  4. For first-party apps, extract operations at build time and enforce trusted documents in production.
  5. Switch hashed queries to GET and set Cache-Control from the shortest-lived field so the CDN can help.
  6. Audit non-null declarations on fields that cross a service boundary.
  7. Alert on the errors array and on per-field error rates, not only on HTTP status.
  8. Adopt federation only when two or more teams need to ship schema changes independently.
Key takeaway: A GraphQL service is a schema plus an engine that calls one resolver per field, so its performance and safety come from what you put around that engine. Batch every lookup with per-request loaders, require pagination and reject queries over a static cost budget, use trusted documents so queries become reviewable and persisted queries so they become cacheable, keep cross-service fields nullable, and read the errors array in monitoring. Reach for federation only when several teams genuinely need to own parts of one graph.