In an event-driven architecture, services communicate by recording facts and letting anyone interested react to them. The accounts service does not call billing, email and provisioning when someone signs up; it records that an account was created, and each of those services decides for itself what that means. The result is loose coupling in time and in knowledge: the producer does not wait for consumers and does not know who they are.

That decoupling is real, and so is its price: eventual consistency, duplicate deliveries, harder debugging and contracts that outlive the teams who wrote them. This article is the decision framework. It separates the different things people mean by event-driven, defines the contracts producers and consumers must honour, and works through a signup flow end to end. The broker mechanics are covered in the publish/subscribe article and a full order system in designing event-driven orders.

Advertisement

Events, commands and queries

Three kinds of message travel between services, and mixing them up causes most event-driven designs to go wrong. A command asks a specific recipient to do something and can be refused: CreateWorkspace. A query asks for data and expects an answer. An event states that something already happened, in the past tense, and cannot be refused: AccountCreated.

The test is ownership of the decision. With a command, the sender decides what should happen and the receiver may say no. With an event, the producer has already decided and committed its own state; consumers decide only their own reaction. If a message called SendWelcomeEmail is published to a topic, it is a command dressed as an event: the producer still owns the decision, it just lost the ability to see whether it succeeded. Name events after facts in the producer's domain, and send commands point to point where someone needs an answer.

The four styles people mean

Event-driven is not one pattern. Martin Fowler's separation into four styles is the most useful map, because each buys something different and costs something different.

StyleWhat the event carriesWhat it buysWhat it costs
Event notificationAn id and a type: account 4821 was createdMinimal coupling; small messagesConsumers call back for details, re-coupling at runtime
Event-carried state transferThe state consumers needConsumers work when the producer is down; no callbacksDuplicated data, larger schemas, eventual consistency
Event sourcingEvery state change, as the system of recordFull history, rebuildable state, auditReplay cost, versioning every event forever
CQRSEvents feed read models separate from the write modelRead models shaped for each queryTwo models to keep in step, lag visible to users

Most systems should start with the second style between services. Notification looks cheaper but every consumer calling back into the producer brings back the synchronous dependency you were trying to remove. Event sourcing and CQRS are internal design choices for one service, not integration styles; see event sourcing before adopting it.

Advertisement

Worked example: account signup

A SaaS product creates accounts. On signup four things must follow: provisioning creates a workspace, email sends a welcome message, billing starts a 14-day trial clock and analytics counts the funnel step. In a synchronous design the accounts service calls all four, and its availability becomes the product of five availabilities. If email is down, signup fails, or worse, half-succeeds.

In the event-driven design the accounts service writes the account and one account.created event in a single database transaction. A relay publishes it to a topic keyed by account id, and each consumer reads at its own pace with its own offset.

Account signup, event-driven: one fact recorded once, consumed independentlyAccounts serviceowns the account rowaccounts DBrow + outbox, one txncommitoutbox relaypolls or reads CDCbroker topicaccount-events, key = account_idpublishProvisioningcreates workspaceEmailsends welcomeBillingstarts trial clockAnalyticsfunnel countsdead-letter topicpoison eventsafter N retriesEach consumer keeps its own offset, its own store and a processed-events table, so a slow or failing consumer delays only itself.The producer never learns who consumed the event; adding a fifth consumer changes no producer code.
The signup flow. The producer commits the account and the event together; four independent consumers react, each idempotently, with a dead-letter path for events that cannot be processed.

The event uses the CloudEvents envelope, whose required attributes are id, source, specversion and type. The version in the type name, and the account_version field, matter later.

{
  "specversion": "1.0",
  "id": "8f0c6a52-6d1e-4b0f-9a57-2f1d2c3b9e11",
  "source": "/accounts",
  "type": "com.example.account.created.v1",
  "subject": "acct_4821",
  "time": "2026-10-02T05:09:14Z",
  "datacontenttype": "application/json",
  "data": {
    "account_id": "acct_4821",
    "plan": "team-trial",
    "owner_email_hash": "sha256:5e88...",
    "region": "eu-west",
    "account_version": 1
  }
}

Note what is absent: the raw email address. Events are copied into many stores with long retention, so carry only what consumers need, and fetch personal data from its owner under access control, or hash it.

The producer contract: commit the fact and the event together

The classic bug is the dual write: update the database, then publish to the broker. If the process dies between the two, the account exists and nobody hears about it; publish first and you can announce an account that was never committed. The fix is the transactional outbox: write the event into an outbox table inside the same transaction, and let a relay publish committed rows, by polling or by change data capture.

def create_account(db, req):
    with db.transaction() as tx:
        acct = tx.insert_account(req.owner, req.plan, req.region)
        tx.execute(
            "INSERT INTO outbox (id, topic, key, type, payload) VALUES (%s, %s, %s, %s, %s)",
            (new_uuid(), "account-events", acct.id,
             "com.example.account.created.v1", to_json(event_for(acct))),
        )
    return acct            # the relay publishes the outbox row after commit

The relay may publish a row twice if it crashes after publishing and before marking the row sent, so the outbox gives at-least-once delivery, never exactly-once. That is fine, because the consumer contract absorbs it. The relay, polling intervals and cleanup are covered in the outbox pattern article.

The consumer contract: idempotent, version-aware, offset last

Every consumer must assume each event can arrive more than once and, across keys, in any order. Three rules make that safe. Record the event id in a processed-events table in the same transaction as the side effect, so a redelivery becomes a no-op. Compare an entity version so an older event cannot overwrite newer state. Commit the broker offset only after the transaction commits.

def on_account_created(event, db):
    with db.transaction() as tx:
        fresh = tx.execute(
            "INSERT INTO processed_events (consumer, event_id) VALUES (%s, %s) "
            "ON CONFLICT DO NOTHING",
            ("billing", event["id"]),
        ).rowcount
        if fresh == 0:
            return "duplicate"                     # redelivery: already applied

        data = event["data"]
        current = tx.get_trial(data["account_id"])
        if current and current.account_version >= data["account_version"]:
            return "stale"                         # older than what we already hold

        tx.upsert_trial(data["account_id"], plan=data["plan"],
                        started_at=event["time"],
                        account_version=data["account_version"])
    return "applied"                               # commit the offset only after this

Side effects outside your database, such as sending an email, cannot share that transaction. Pass the event id to the downstream API as an idempotency key, or accept that a crash between send and commit will occasionally send twice, and decide whether that is tolerable. That decision belongs in the design document, not in a surprise incident.

Keys, partitions and the order you can rely on

Log brokers such as Kafka guarantee order only within a partition, and the producer's key chooses the partition. Key every event about an account by account id and all of that account's events arrive at a consumer in the order they were published. Nothing orders events about different accounts, and nothing should need to.

Order is lost in three common ways. Retrying a failed event by republishing it to the same topic puts it behind later events for the same key. Processing a partition with a thread pool, without keeping per-key order, reorders silently. And a producer with several writers per entity, with no single owner, can publish versions 2 and 3 in the wrong order. The version check in the consumer is the safety net for all three; a single writer per entity and per-key serial processing are the prevention.

Choreography or orchestration

Signup is choreography: each consumer reacts independently, and there is no workflow to coordinate, because none of the four reactions depends on another. That is where choreography is at its best. It is at its worst when a business process has steps that must happen in order and be compensated on failure, because the process then exists only implicitly, spread over the subscriptions of several services, and nobody can answer where a given signup is stuck.

A practical rule: choreograph independent reactions to a fact; orchestrate a process that has an owner, an order and a deadline. If converting a trial to a paid plan requires charging a card, upgrading limits and sending a receipt, with a refund if the upgrade fails, write an orchestrator that issues commands and listens for the resulting events. It can still be event-driven underneath, but the process state lives in one place you can query.

Evolving event schemas without breaking consumers

An event schema is a public API whose consumers you may not know. Additive changes are safe if consumers follow the tolerant reader rule: ignore fields you do not recognise, and never fail on a missing optional field. Renaming a field, changing its type or changing its meaning is a breaking change.

For a breaking change, publish a new event type, such as account.created.v2, alongside the old one, migrate consumers, and retire v1 when its consumer lag and traffic show nobody reads it. A schema registry with a compatibility mode, typically backward compatibility so new readers handle old data, turns accidental breaking changes into failed producer builds instead of failed consumers. Remember that retained events are replayed: a consumer must still read every version still held on the topic.

Failure modes

SymptomCauseFix
Account exists, no welcome emailDual write; publish lost after commitTransactional outbox
Trial started twiceRedelivery with a non-idempotent handlerProcessed-events table in the same transaction
Newer plan overwritten by older oneOut-of-order events for one keyKey by entity id; version check in consumer
One bad event blocks a partitionPoison message retried foreverBounded retries, then dead-letter topic with alerting
Consumer falls hours behindSlow handler or undersized consumer groupAlert on lag; scale up to the partition count
Replay sends emails againSide effects not guarded during reprocessingIdempotency keys downstream; a replay mode that suppresses effects

Operating it

Measure each consumer group's lag, in messages and in seconds since the oldest unprocessed event; seconds is what users feel. Track end-to-end latency from the event's time to the consumer's commit, dead-letter depth and age, and outbox rows older than a minute, which means the relay is stuck. Propagate the W3C traceparent header in event headers so a trace can follow a signup from the HTTP request through every consumer; without it, debugging is grep across four log streams.

Run a replay drill before you need one: reset one consumer group to an earlier offset in a staging environment and confirm that nothing user-visible happens twice. Document each topic's retention, owner, key and schema, and the consumers you know about.

When not to go event-driven

Use a synchronous call when the caller needs the answer to continue, such as checking whether a username is taken, or when the user must see the effect immediately and cannot tolerate lag. Keep a small system with one team and one database as a modular monolith; events across modules inside one process add a broker, retries and lag for no decoupling benefit. Event-driven pays when several teams react independently to the same facts, when producers must survive consumer outages, and when traffic is bursty enough that a buffer between them is valuable.

What to do next

  1. List the messages your services exchange and label each as command, query or event; rename or reroute any command disguised as an event.
  2. For each event, choose notification or state transfer deliberately, and remove personal data consumers do not need.
  3. Replace every database-then-publish dual write with an outbox, and make every consumer idempotent with a processed-events table and a version check.
  4. Key events by entity id and keep per-key processing serial.
  5. Put event schemas in a registry with a compatibility mode, and adopt versioned event types for breaking changes.
  6. Add lag, end-to-end latency, dead-letter and outbox-age alerts, and propagate trace context in event headers. Then read the idempotency article for the downstream side effects.
Key takeaway: Event-driven architecture means producers record facts and consumers decide their own reactions, which removes synchronous coupling at the price of eventual consistency and duplicate delivery. Prefer event-carried state transfer between services, commit each fact and its event in one transaction with an outbox, and make every consumer idempotent and version-aware. Choreograph independent reactions, orchestrate processes that have an owner and an order, evolve schemas additively, and watch lag in seconds.