In an event-driven architecture, services communicate by recording facts and letting anyone interested react to them. The accounts service does not call billing, email and provisioning when someone signs up; it records that an account was created, and each of those services decides for itself what that means. The result is loose coupling in time and in knowledge: the producer does not wait for consumers and does not know who they are.
That decoupling is real, and so is its price: eventual consistency, duplicate deliveries, harder debugging and contracts that outlive the teams who wrote them. This article is the decision framework. It separates the different things people mean by event-driven, defines the contracts producers and consumers must honour, and works through a signup flow end to end. The broker mechanics are covered in the publish/subscribe article and a full order system in designing event-driven orders.
Events, commands and queries
Three kinds of message travel between services, and mixing them up causes most event-driven designs to go wrong. A command asks a specific recipient to do something and can be refused: CreateWorkspace. A query asks for data and expects an answer. An event states that something already happened, in the past tense, and cannot be refused: AccountCreated.
The test is ownership of the decision. With a command, the sender decides what should happen and the receiver may say no. With an event, the producer has already decided and committed its own state; consumers decide only their own reaction. If a message called SendWelcomeEmail is published to a topic, it is a command dressed as an event: the producer still owns the decision, it just lost the ability to see whether it succeeded. Name events after facts in the producer's domain, and send commands point to point where someone needs an answer.
The four styles people mean
Event-driven is not one pattern. Martin Fowler's separation into four styles is the most useful map, because each buys something different and costs something different.
| Style | What the event carries | What it buys | What it costs |
|---|---|---|---|
| Event notification | An id and a type: account 4821 was created | Minimal coupling; small messages | Consumers call back for details, re-coupling at runtime |
| Event-carried state transfer | The state consumers need | Consumers work when the producer is down; no callbacks | Duplicated data, larger schemas, eventual consistency |
| Event sourcing | Every state change, as the system of record | Full history, rebuildable state, audit | Replay cost, versioning every event forever |
| CQRS | Events feed read models separate from the write model | Read models shaped for each query | Two models to keep in step, lag visible to users |
Most systems should start with the second style between services. Notification looks cheaper but every consumer calling back into the producer brings back the synchronous dependency you were trying to remove. Event sourcing and CQRS are internal design choices for one service, not integration styles; see event sourcing before adopting it.
Worked example: account signup
A SaaS product creates accounts. On signup four things must follow: provisioning creates a workspace, email sends a welcome message, billing starts a 14-day trial clock and analytics counts the funnel step. In a synchronous design the accounts service calls all four, and its availability becomes the product of five availabilities. If email is down, signup fails, or worse, half-succeeds.
In the event-driven design the accounts service writes the account and one account.created event in a single database transaction. A relay publishes it to a topic keyed by account id, and each consumer reads at its own pace with its own offset.
The event uses the CloudEvents envelope, whose required attributes are id, source, specversion and type. The version in the type name, and the account_version field, matter later.
{
"specversion": "1.0",
"id": "8f0c6a52-6d1e-4b0f-9a57-2f1d2c3b9e11",
"source": "/accounts",
"type": "com.example.account.created.v1",
"subject": "acct_4821",
"time": "2026-10-02T05:09:14Z",
"datacontenttype": "application/json",
"data": {
"account_id": "acct_4821",
"plan": "team-trial",
"owner_email_hash": "sha256:5e88...",
"region": "eu-west",
"account_version": 1
}
}Note what is absent: the raw email address. Events are copied into many stores with long retention, so carry only what consumers need, and fetch personal data from its owner under access control, or hash it.
The producer contract: commit the fact and the event together
The classic bug is the dual write: update the database, then publish to the broker. If the process dies between the two, the account exists and nobody hears about it; publish first and you can announce an account that was never committed. The fix is the transactional outbox: write the event into an outbox table inside the same transaction, and let a relay publish committed rows, by polling or by change data capture.
def create_account(db, req):
with db.transaction() as tx:
acct = tx.insert_account(req.owner, req.plan, req.region)
tx.execute(
"INSERT INTO outbox (id, topic, key, type, payload) VALUES (%s, %s, %s, %s, %s)",
(new_uuid(), "account-events", acct.id,
"com.example.account.created.v1", to_json(event_for(acct))),
)
return acct # the relay publishes the outbox row after commitThe relay may publish a row twice if it crashes after publishing and before marking the row sent, so the outbox gives at-least-once delivery, never exactly-once. That is fine, because the consumer contract absorbs it. The relay, polling intervals and cleanup are covered in the outbox pattern article.
The consumer contract: idempotent, version-aware, offset last
Every consumer must assume each event can arrive more than once and, across keys, in any order. Three rules make that safe. Record the event id in a processed-events table in the same transaction as the side effect, so a redelivery becomes a no-op. Compare an entity version so an older event cannot overwrite newer state. Commit the broker offset only after the transaction commits.
def on_account_created(event, db):
with db.transaction() as tx:
fresh = tx.execute(
"INSERT INTO processed_events (consumer, event_id) VALUES (%s, %s) "
"ON CONFLICT DO NOTHING",
("billing", event["id"]),
).rowcount
if fresh == 0:
return "duplicate" # redelivery: already applied
data = event["data"]
current = tx.get_trial(data["account_id"])
if current and current.account_version >= data["account_version"]:
return "stale" # older than what we already hold
tx.upsert_trial(data["account_id"], plan=data["plan"],
started_at=event["time"],
account_version=data["account_version"])
return "applied" # commit the offset only after thisSide effects outside your database, such as sending an email, cannot share that transaction. Pass the event id to the downstream API as an idempotency key, or accept that a crash between send and commit will occasionally send twice, and decide whether that is tolerable. That decision belongs in the design document, not in a surprise incident.
Keys, partitions and the order you can rely on
Log brokers such as Kafka guarantee order only within a partition, and the producer's key chooses the partition. Key every event about an account by account id and all of that account's events arrive at a consumer in the order they were published. Nothing orders events about different accounts, and nothing should need to.
Order is lost in three common ways. Retrying a failed event by republishing it to the same topic puts it behind later events for the same key. Processing a partition with a thread pool, without keeping per-key order, reorders silently. And a producer with several writers per entity, with no single owner, can publish versions 2 and 3 in the wrong order. The version check in the consumer is the safety net for all three; a single writer per entity and per-key serial processing are the prevention.
Choreography or orchestration
Signup is choreography: each consumer reacts independently, and there is no workflow to coordinate, because none of the four reactions depends on another. That is where choreography is at its best. It is at its worst when a business process has steps that must happen in order and be compensated on failure, because the process then exists only implicitly, spread over the subscriptions of several services, and nobody can answer where a given signup is stuck.
A practical rule: choreograph independent reactions to a fact; orchestrate a process that has an owner, an order and a deadline. If converting a trial to a paid plan requires charging a card, upgrading limits and sending a receipt, with a refund if the upgrade fails, write an orchestrator that issues commands and listens for the resulting events. It can still be event-driven underneath, but the process state lives in one place you can query.
Evolving event schemas without breaking consumers
An event schema is a public API whose consumers you may not know. Additive changes are safe if consumers follow the tolerant reader rule: ignore fields you do not recognise, and never fail on a missing optional field. Renaming a field, changing its type or changing its meaning is a breaking change.
For a breaking change, publish a new event type, such as account.created.v2, alongside the old one, migrate consumers, and retire v1 when its consumer lag and traffic show nobody reads it. A schema registry with a compatibility mode, typically backward compatibility so new readers handle old data, turns accidental breaking changes into failed producer builds instead of failed consumers. Remember that retained events are replayed: a consumer must still read every version still held on the topic.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Account exists, no welcome email | Dual write; publish lost after commit | Transactional outbox |
| Trial started twice | Redelivery with a non-idempotent handler | Processed-events table in the same transaction |
| Newer plan overwritten by older one | Out-of-order events for one key | Key by entity id; version check in consumer |
| One bad event blocks a partition | Poison message retried forever | Bounded retries, then dead-letter topic with alerting |
| Consumer falls hours behind | Slow handler or undersized consumer group | Alert on lag; scale up to the partition count |
| Replay sends emails again | Side effects not guarded during reprocessing | Idempotency keys downstream; a replay mode that suppresses effects |
Operating it
Measure each consumer group's lag, in messages and in seconds since the oldest unprocessed event; seconds is what users feel. Track end-to-end latency from the event's time to the consumer's commit, dead-letter depth and age, and outbox rows older than a minute, which means the relay is stuck. Propagate the W3C traceparent header in event headers so a trace can follow a signup from the HTTP request through every consumer; without it, debugging is grep across four log streams.
Run a replay drill before you need one: reset one consumer group to an earlier offset in a staging environment and confirm that nothing user-visible happens twice. Document each topic's retention, owner, key and schema, and the consumers you know about.
When not to go event-driven
Use a synchronous call when the caller needs the answer to continue, such as checking whether a username is taken, or when the user must see the effect immediately and cannot tolerate lag. Keep a small system with one team and one database as a modular monolith; events across modules inside one process add a broker, retries and lag for no decoupling benefit. Event-driven pays when several teams react independently to the same facts, when producers must survive consumer outages, and when traffic is bursty enough that a buffer between them is valuable.
What to do next
- List the messages your services exchange and label each as command, query or event; rename or reroute any command disguised as an event.
- For each event, choose notification or state transfer deliberately, and remove personal data consumers do not need.
- Replace every database-then-publish dual write with an outbox, and make every consumer idempotent with a processed-events table and a version check.
- Key events by entity id and keep per-key processing serial.
- Put event schemas in a registry with a compatibility mode, and adopt versioned event types for breaking changes.
- Add lag, end-to-end latency, dead-letter and outbox-age alerts, and propagate trace context in event headers. Then read the idempotency article for the downstream side effects.