An API is a promise you make to code you will never see. Once a client depends on a field, a status code or even the order of keys in a response, changing it breaks someone, and the observation that every observable behaviour eventually becomes depended upon (often called Hyrum's law) applies with full force. Good API design is therefore mostly about deciding, before the first client arrives, which behaviours you are willing to promise forever and which you keep private.
This guide walks the design process end to end on one worked example, an orders API for an online bookshop. It covers the decisions that are expensive to reverse: resource names and identifiers, the error model, retries and idempotency, pagination, concurrent updates, rate limits and versioning. The examples are HTTP and JSON because that is what most public APIs use, but the reasoning carries over to gRPC, and a section at the end compares the styles.
Start from the consumer, not the database
The most common design mistake is exposing your tables. A table layout reflects how you store data today; an API should reflect what callers are trying to do. Before drawing a single endpoint, write down the use cases as sentences with a caller, a goal and a frequency: the storefront creates an order from a cart, about 30 times a second at peak; the warehouse lists orders ready to ship, every minute; a customer cancels an unshipped order; finance exports a day of orders.
Each sentence implies a requirement. The storefront needs creation to be safe to retry. The warehouse needs a stable list it can page through while new orders arrive. Cancellation is a state transition, not an edit. Finance needs bulk access, which belongs in an export job rather than a paginated endpoint called 50,000 times in a loop.
Write the contract before the implementation. With OpenAPI for HTTP or a .proto file for gRPC, the contract is a reviewable artifact: consumers can comment on it, a mock server can be generated so they start building, and a linter can enforce your house rules. The diagram shows the flow; the key property is that the implementation is the last thing to exist, so the design is not bent around whatever was easiest to code.
Model resources and their states
Resources are nouns with identity and a lifecycle. For the bookshop the core resources are orders, the line items inside an order, and shipments. Write the order lifecycle as a state machine before naming URLs: pending_payment to paid to shipped to delivered, with cancelled reachable only from the first two. Transitions that carry business meaning become explicit operations; fields that callers may freely edit become plain updates.
| Operation | Request | Notes |
|---|---|---|
| Create order | POST /v1/orders | Requires an Idempotency-Key header; returns 201 with Location |
| Get order | GET /v1/orders/{order_id} | Returns an ETag for concurrency control |
| List orders | GET /v1/orders?status=paid&page_size=50 | Cursor pagination, stable sort by creation time then id |
| Update address | PATCH /v1/orders/{order_id} | Only shipping_address is mutable; requires If-Match |
| Cancel order | POST /v1/orders/{order_id}:cancel | State transition; 409 if already shipped |
| List shipments | GET /v1/orders/{order_id}/shipments | Child collection; one order may ship in parts |
A few rules make this consistent. Use plural nouns for collections and opaque string identifiers such as ord_8f3k2m rather than auto-increment integers, which leak volume and invite enumeration. Prefix identifiers with a type so a misplaced id fails loudly. Represent money as an integer amount in minor units plus an ISO 4217 currency code, never as a float. Use RFC 3339 timestamps in UTC. Keep field names in one case convention across the whole API.
Custom verbs such as :cancel are a deliberate choice; some teams prefer POST /v1/orders/{id}/cancellation. Either is fine if it is used consistently. What is not fine is cancelling by PATCHing status directly, because then every client can attempt any transition and the server must reverse-engineer intent from a field diff.
An error model clients can program against
Clients branch on errors, so errors are part of the contract. Status codes give the coarse class: 400 for a malformed request, 401 for missing or invalid credentials, 403 for an authenticated caller without permission, 404 for an unknown resource, 409 for a conflict with current state, 412 for a failed precondition, 422 when the request is well formed but semantically invalid, 429 for rate limiting, and 5xx only for server faults. A body is needed for the detail. RFC 9457, which obsoleted RFC 7807, defines the application/problem+json format with type, title, status, detail and instance members, and allows extension members.
HTTP/1.1 409 Conflict
Content-Type: application/problem+json
{
"type": "https://api.example-books.com/problems/order-not-cancellable",
"title": "Order cannot be cancelled",
"status": 409,
"detail": "Order ord_8f3k2m shipped at 2026-09-30T14:02:11Z",
"instance": "/v1/orders/ord_8f3k2m:cancel",
"order_state": "shipped",
"request_id": "req_01J9ZK4"
}The type URI is the stable, machine-readable code that clients switch on; title and detail are for humans and may change wording. Validation errors should list every invalid field at once in an extension member. Include a request id in every response, success or failure, so a support ticket can be joined to server logs. Never put stack traces, SQL or internal hostnames in an error body.
Retries and idempotency
Networks fail after the server has done the work but before the client hears about it. GET, PUT and DELETE are defined as idempotent by HTTP semantics, so a client may simply retry them. POST is not: a retried create can make two orders and charge twice. The widely used fix is an idempotency key: the client generates a unique key per logical operation and sends it in a header, and the server stores the first result under that key and replays it for any retry. An IETF draft standardises the Idempotency-Key header name; it is still a draft, but the pattern is established practice.
def create_order(request, db):
key = request.headers.get("Idempotency-Key")
if not key:
return problem(400, "idempotency-key-required")
fingerprint = sha256(canonical_json(request.body))
with db.transaction():
row = db.select_for_update("idem", (request.client_id, key))
if row:
if row.fingerprint != fingerprint:
return problem(422, "idempotency-key-reused") # same key, different body
if row.state == "in_progress":
return problem(409, "request-in-progress") # concurrent duplicate
return replay(row.status, row.body) # stored first result
db.insert("idem", (request.client_id, key), fingerprint, state="in_progress")
order = place_order(request.body) # same transaction if possible; else expire stale in_progress rows
db.update("idem", (request.client_id, key), state="done", status=201, body=order)
return respond(201, order)Three details matter. Scope keys per client so two tenants cannot collide. Store a fingerprint of the request body and reject a reused key with a different body, because that is a client bug, not a retry. Expire keys after a documented window, for example 24 hours, and say so in the docs. The dedupe store design, including what happens when the work and the key live in different databases, is covered in idempotency architecture.
Pagination that survives concurrent writes
Offset pagination (?offset=200&limit=50) is simple but breaks under writes: an insert near the front shifts every later page, so clients skip or repeat rows, and large offsets get slower because the database must walk past them. Cursor pagination fixes both. Sort by a unique, immutable key such as (created_at, id), and return an opaque cursor that encodes the last key seen.
import base64, json
def encode_cursor(last_row, filters):
raw = json.dumps({"t": last_row.created_at.isoformat(), "id": last_row.id, "f": filters})
return base64.urlsafe_b64encode(raw.encode()).decode() # sign or encrypt if clients must not edit it
def list_orders(db, filters, page_size=50, cursor=None):
page_size = min(page_size, 200) # server-enforced maximum
after = json.loads(base64.urlsafe_b64decode(cursor)) if cursor else None
if after and after["f"] != filters:
raise BadRequest("cursor does not match filters")
rows = db.query(
"SELECT * FROM orders WHERE status = %s"
" AND (%s IS NULL OR (created_at, id) > (%s, %s))"
" ORDER BY created_at, id LIMIT %s",
filters["status"], after and after["t"], after and after["t"], after and after["id"], page_size + 1)
more = len(rows) > page_size
rows = rows[:page_size]
return {"data": rows, "next_cursor": encode_cursor(rows[-1], filters) if more else None}Fetching one extra row tells you whether another page exists without an expensive count query. Document the cursor as opaque so you can change its contents later, and do not promise a total count unless a use case needs it.
Concurrent updates: ETags and If-Match
Two support agents open the same order and both edit the address. Without protection the second save silently overwrites the first. Optimistic concurrency solves this without locks: every representation carries a version, exposed as an ETag header, and updates must send it back in If-Match. The server performs the update only if the stored version still matches, otherwise it returns 412 Precondition Failed and the client re-reads.
UPDATE orders
SET shipping_address = $1, version = version + 1
WHERE id = $2 AND version = $3 -- $3 comes from If-Match
RETURNING version;
-- zero rows updated: 412 if the order exists, 404 if it does notYou can make If-Match mandatory on PATCH and answer 428 Precondition Required when it is absent; that forces every client to handle conflicts rather than discovering them in production.
Evolution, versioning and deprecation
Most change should be additive. Adding an optional request field, a new response field, a new endpoint or a new enum value that clients were told to tolerate are compatible changes. Removing or renaming a field, changing a type, tightening validation, changing a default, or changing the meaning of an existing value are breaking. Write the rule into the contract docs: clients must ignore unknown fields and must handle unknown enum values with a fallback. Then enforce it on the server side with a compatibility diff in CI that fails the build when the new OpenAPI or proto file breaks the old one.
When a break is unavoidable, introduce a new major version, usually in the path (/v2/) because it is visible in logs, caches and support tickets. Header- or date-based versioning lets you evolve finer-grained, at the cost of more complex routing and testing. Run both versions, measure who still calls the old one per client id, and announce removal with the Sunset response header (RFC 8594) plus the Deprecation header, standardised as RFC 9745. Then contact the remaining callers directly.
Security, limits and the cross-cutting layer
Authenticate with short-lived bearer tokens, for example OAuth 2.0 access tokens with scopes such as orders:read and orders:write, and authorize every request against the resource owner, not just the scope; the classic API vulnerability is an endpoint that checks the caller is logged in but not that order ord_8f3k2m belongs to them. Never accept credentials in query strings, which end up in logs and browser history.
Rate limits protect shared capacity. Return 429 with a Retry-After header, document the limits per plan, and expect clients to back off with jitter. The IETF RateLimit header fields that advertise remaining quota are still a draft, so if you emit them, say which draft revision you follow. Limiting algorithms are covered in rate limiter design.
REST, gRPC or GraphQL
| Style | Strengths | Costs | Fits |
|---|---|---|---|
| Resource-oriented HTTP + JSON | Universal tooling, cacheable GETs, easy to debug with curl | Verbose payloads, no built-in streaming contract, schema discipline is optional | Public and partner APIs |
| gRPC | Strong schema, generated clients, efficient binary encoding, bidirectional streaming | Needs HTTP/2 end to end, harder from browsers, opaque on the wire | Internal service-to-service traffic |
| GraphQL | Clients choose fields, one round trip for nested data | Query cost control, caching and authorization per field are hard | Product front ends aggregating many backends |
The design rules in this guide apply to all three: stable identifiers, explicit state transitions, typed errors, idempotent writes and additive evolution. For gRPC specifics see gRPC architecture.
Failure modes
| Failure | What goes wrong | Defence |
|---|---|---|
| Leaky data model | Every schema migration becomes an API break | Design from use cases; map storage to resources in one layer |
| Untyped errors | Clients parse message strings that later change | Stable problem type URIs; human text is free to change |
| Non-idempotent create | Duplicate orders and double charges on retry | Idempotency keys with fingerprint and stored result |
| Offset pagination on hot tables | Skipped and repeated rows, slow deep pages | Keyset cursor on a unique sort key |
| Lost updates | Last writer silently wins | ETag with mandatory If-Match, 412 on mismatch |
| Broken object-level authorization | Callers read others' orders by guessing ids | Ownership check on every request; opaque ids as defence in depth |
| Silent breaking change | A deploy breaks clients nobody knew existed | Compatibility diff in CI, per-client usage metrics |
Worked review of the bookshop design
Run the use cases back through the design. The storefront retries a timed-out create with the same key and gets the original 201 replayed, so one order and one charge. The warehouse pages status=paid by cursor; orders created mid-scan appear on a later page instead of shifting earlier ones. A customer cancelling after shipment gets a 409 with a typed problem the app can turn into a message offering a return. Two agents editing an address get one success and one 412. Finance does not hammer the list endpoint: an asynchronous export job writes a file and returns a link. Every one of those outcomes was decided in the contract, which is the point. For writing the design up for review, see design docs.
What to do next
- Write five to ten use-case sentences with caller, goal and frequency before naming any endpoint.
- Draw the state machine for each resource and turn business transitions into explicit operations.
- Write the OpenAPI or proto contract first, generate a mock server and get a consumer to build against it.
- Adopt problem+json errors with stable type URIs and a request id on every response.
- Require Idempotency-Key on every non-idempotent create and document the retention window.
- Use keyset cursors with a server-enforced maximum page size; drop total counts unless a use case needs them.
- Return ETags and require If-Match on updates.
- Add a compatibility diff to CI and record per-client, per-version usage so you know who a change will break.