Microservices architecture splits one application into several services. Each one is deployed on its own, owns its own data and talks to the others only over the network through published contracts. The name says small, but size is not the point. The point is independence: a team can change, test, release and scale its service without coordinating a release with every other team. Everything else in the style, including its considerable cost, follows from that goal.

This article explains the architecture as a whole. It covers how to draw service boundaries, why a database per service is the rule that matters most, when to call another service synchronously and when to publish an event, and the arithmetic that makes a deep call graph fragile. It then walks through a checkout across six services, the platform you need before the second service ships, the common failure modes, and the honest case for staying with a modular monolith. Detailed mechanisms such as sagas, the outbox, gateways and service meshes have their own pages, linked where they come up.

Advertisement

What makes a service a microservice

A useful definition has three tests. First, the service is independently deployable: you can release it without releasing anything else, and its consumers keep working. Second, it owns its data: no other service reads or writes its tables, and the only way to reach the data is through its API or its events. Third, it is owned by one team, which is on call for it and decides its roadmap. A codebase can be split into twenty deployables and still fail all three tests. If every release needs a coordinated train, or four services share one schema, you have a distributed monolith. That is the worst of both styles: network failures without independence.

The style was a reaction to large teams blocking each other on a single deployable. It trades in-process function calls, which are fast, reliable and transactional, for network calls, which are slow, can fail partially and cannot share a transaction. That trade only pays when the organisational gain is real. Conway's law cuts both ways here. Systems mirror the communication structure of the teams that build them, so service boundaries that do not match team boundaries produce constant cross-team negotiation.

Drawing boundaries around data and behaviour

Start from the business capabilities and the data each one is the authority for, not from technical layers. A user service, a database service and an email service is a layered split, and almost every feature will touch all three. Ordering, payments, inventory, pricing and notifications are capabilities. Each one owns a clear set of facts: the order and its status, the payment and its capture, stock levels and reservations. Domain-driven design calls these bounded contexts. Inside one context a word such as product has a single meaning. At the edge it is translated, because to inventory a product is a stock-keeping unit with a bin location, and to pricing it is a price list entry.

A good boundary has high cohesion inside and narrow, stable contracts outside. One test is to list the last twenty feature requests and count how many services each would have touched. If most touch three or more, the boundaries are wrong. Another test is the transaction: if two pieces of data must change atomically in nearly every operation, they belong in the same service. Splitting them forces a saga for something a local transaction did for free.

Database per service is the rule that protects independence. It does not have to mean a separate database server; separate schemas with separate credentials work. What matters is that no service can query another's tables. Shared tables are a hidden contract that nobody versions. The first time inventory renames a column, orders breaks in production.

Advertisement

Reference architecture

Checkout as microservices: each service owns its data; events cross boundariesclientsweb, mobileAPI gatewayauth, routing, limitsOrder serviceorders DB + outboxsyncPricing serviceread-only, cachedsync, 150 msevent brokerOrderPlaced, PaymentCaptured, StockReservedrelayPayment servicepayments DBInventory servicestock DBNotification serviceno shared tablesSearch read modelprojectionplatformCI/CD per service, service discovery, tracing, metrics, logs, secrets, mTLSOnly the gateway-to-order and order-to-pricing hops are synchronous; everything else reacts to events.
A checkout split into services. Clients enter through a gateway, the order service makes one bounded synchronous call for pricing, and the rest of the flow is driven by events published through a transactional outbox. Read models such as search are projections built from events.

Notice how few synchronous edges there are. The gateway handles authentication, routing and rate limits (see API gateway architecture). Instances find each other through service discovery, and many platforms put mutual TLS, retries and telemetry into a service mesh so each service does not reimplement them.

Synchronous calls versus events

Every interaction between services is either a request that waits for an answer or a message that is published and handled later. The choice decides coupling in time. A synchronous call means the caller is only up when the callee is up. An event means the producer does not care whether the consumer is running right now.

QuestionSynchronous request (HTTP, gRPC)Asynchronous event (broker, log)
Caller needs the answer to continue?Yes: a price, an authorisation decisionNo: tell others that something happened
Availability couplingCaller fails when callee failsConsumer can be down; messages wait
LatencyAdds callee latency to the user's requestOff the request path; eventual
ConsistencyReads the latest stateConsumers see state after a lag
Failure handlingTimeouts, retries, fallbacksRedelivery, idempotent consumers, dead-letter queues
DebuggingOne trace, easy to followNeeds correlation IDs across hops

A practical rule is to make queries synchronous and keep the chain short, and to make state changes that other services react to into events. Commands that must succeed across services, such as reserve stock and then charge the card, become a saga: a sequence of local transactions with compensating actions, covered in the saga pattern.

Failure arithmetic: why deep call graphs break

Availabilities multiply along a synchronous chain. If a request passes through ten services in series and each is available 99.9 percent of the time, the request succeeds about 0.999 to the power 10, roughly 99.0 percent of the time. That is ten times the downtime of any single service. At 99.9 percent, a month allows about 43 minutes of failure; at 99.0 percent it allows about 7.3 hours.

Latency also compounds, and the tail compounds most. If a request fans out to five services in parallel and waits for all of them, its latency is the slowest of the five. With five independent dependencies, the chance that at least one is slower than its own 99th percentile is 1 minus 0.99 to the power 5, about 4.9 percent. Your median request now pays the tail of something.

Retries amplify load. If three layers each make one call plus up to two retries, a single failing leaf can receive 3 times 3 times 3, which is 27 attempts for one user click. That is how a slow database turns into an outage. The fixes are simple to state. Retry at one layer only, usually the one nearest the failure. Cap retries with a budget, for example retries may not exceed 10 percent of requests. Pass a deadline down the chain so a callee stops working once the caller has given up.

// Order service calling Pricing: every remote call gets a deadline, a bounded retry
// and a fallback decision made up front, not after the first outage.
public Price quote(String sku, Duration remaining) {
    Duration budget = min(remaining.minusMillis(50), Duration.ofMillis(150)); // leave time to answer
    if (budget.isNegative()) throw new DeadlineExceeded("no time left for pricing");
    for (int attempt = 1; attempt <= 2; attempt++) {           // at most one retry
        if (!retryBudget.tryAcquire(attempt)) break;           // fleet-wide cap, e.g. 10% of calls
        try {
            return pricing.get("/prices/" + sku)
                          .timeout(budget)
                          .header("x-request-id", requestId())
                          .execute(Price.class);
        } catch (TimeoutException | ServiceUnavailable e) {
            metrics.counter("pricing.failure", "attempt", attempt).increment();
            sleep(jitter(Duration.ofMillis(20 * attempt)));
        }
    }
    return priceCache.lastKnown(sku)                            // explicit, product-approved fallback
                     .orElseThrow(() -> new DependencyDown("pricing"));
}

The code shows the minimum every remote call should carry: a timeout derived from the caller's remaining deadline, one retry with jitter that draws from a shared retry budget, a request ID for tracing and a fallback the product owner agreed to. Breaking the circuit after repeated failures, so callers fail fast instead of queueing, is the next layer.

Worked example: placing an order

Follow one checkout through the diagram. The client sends POST /orders with an Idempotency-Key header. The gateway authenticates the user and routes to the order service. The order service calls pricing synchronously with a 150 millisecond budget. If pricing is down, it uses the last cached price only if the product team has agreed that is acceptable; otherwise it returns 503 and the client retries with the same key.

The order service then writes the order and an OrderPlaced event in one local transaction, which is the transactional outbox described in the outbox pattern. It returns 202 Accepted with the order ID and status PENDING. The user sees an order page at once, and no payment or inventory call sits on the request path.

-- One local transaction: the order row and the event row commit together or not at all.
BEGIN;
WITH ins AS (
  INSERT INTO orders (id, customer_id, total_cents, status, idempotency_key)
  VALUES ($1, $2, $3, 'PENDING', $4)
  ON CONFLICT (idempotency_key) DO NOTHING      -- a retried POST creates nothing new
  RETURNING id
)
INSERT INTO outbox (id, aggregate_id, type, payload, created_at)
SELECT gen_random_uuid(), id, 'OrderPlaced', $5, now() FROM ins;  -- no order row, no event
COMMIT;
-- A relay process reads unsent outbox rows in order, publishes them, then marks them sent.
-- Consumers must tolerate duplicates: the relay can crash after publishing, before marking.

A relay publishes OrderPlaced. Inventory reserves stock and publishes StockReserved or StockUnavailable. Payment, which waits for the reservation, captures the card and publishes PaymentCaptured or PaymentFailed. The order service listens to both and moves the order to CONFIRMED, or to CANCELLED with a compensating StockReleased if payment failed. Notification sends the email when the order is confirmed. Search updates its projection. Every consumer records the event IDs it has processed, because the broker and the relay both deliver at least once.

Count what this design buys. With payment down for ten minutes, users can still place orders; they stay PENDING and are confirmed when payment recovers. With notification down, nothing user-facing breaks. The price is that the order page must show an honest pending state, and support staff need a tool to see where any order is stuck.

Operating many services

A monolith needs one pipeline, one dashboard and one log search. Twenty services need the same things twenty times, with no manual steps. Build this platform before the second service ships, not after the tenth.

  • Independent pipelines. Each service builds, tests and deploys on its own, with canary or blue-green releases and automatic rollback on error-rate or latency regressions.
  • Contract tests. Consumer-driven contract tests run in the provider's pipeline, so a provider change that breaks a consumer fails before release. API changes are additive; removals need a deprecation window.
  • Distributed tracing. Every request and event carries a trace context, for example W3C traceparent, propagated through HTTP headers and message headers.
  • Golden signals per service. Request rate, errors, latency percentiles and saturation, plus consumer lag for every subscription, with service level objectives the owning team is paged for.

Failure modes

SymptomLikely causeFix
Every release needs several teamsShared database or shared domain libraryGive each service its own schema; publish events instead
One slow service takes the site downNo timeouts, unbounded retries, threads exhaustedDeadlines, retry budget, bulkheads, circuit breakers
Request latency grows with every featureLong synchronous call chainsMove reactions to events; cache read-mostly data near callers
Duplicate charges or emailsAt-least-once delivery without idempotencyIdempotency keys on commands; processed-ID table on consumers
Order stuck in PENDING foreverLost event or consumer stuck on a poison messageOutbox, dead-letter queue, alert on age of oldest pending item
Nobody can explain a failureNo correlation across hopsTrace context on every call and message; structured logs with IDs

Trade-offs and the modular monolith

Microservices buy team autonomy, independent scaling and fault isolation. They cost network latency, partial failure, eventual consistency, duplicated platform work and much harder debugging. For a single team, or a product still discovering its domain, the costs usually win. Boundaries drawn too early are expensive to move, because moving one means migrating data between databases rather than moving code between packages.

A modular monolith is the usual alternative. It is one deployable with strict internal modules that own their own tables and talk through interfaces, with the rules enforced by build-time dependency checks. You keep local transactions and simple operations while you learn where the real seams are. When one module needs a separate release cadence, a different scaling profile or a separate team, extract it with the strangler fig approach. Route its traffic through a facade, build the new service behind it, migrate the data, then cut over route by route.

What to do next

  1. Write down, for each existing or planned service, the data it is the authority for and the team that owns it; merge any services that share tables or a team.
  2. Draw the synchronous call graph for your top three user requests, compute the chained availability, and convert any edge that is a notification rather than a query into an event.
  3. Give every remote call a deadline, at most one retry layer, a retry budget and a documented fallback.
  4. Put state changes that others react to behind a transactional outbox, and make every consumer idempotent with a processed-ID table.
  5. Add trace propagation through both HTTP and message headers, and alert on the age of the oldest unprocessed message per consumer.
  6. If you have one team, build a modular monolith with enforced module boundaries first, and extract a service only for a named release, scaling or compliance reason.
Key takeaway: A microservice is independently deployable, owns its data and is owned by one team; without all three you have a distributed monolith. Draw boundaries around business capabilities and the facts each one is authoritative for, keep synchronous chains short because availability multiplies and retries amplify, and let state changes travel as events through an outbox to idempotent consumers. Build the delivery, tracing and contract-testing platform before the second service, and prefer a modular monolith until a team, scaling or compliance need justifies the network.