A saga replaces one distributed transaction with a sequence of local transactions, each with a compensating action that undoes it if a later step fails. The mechanics of compensation, the pivot step and isolation are covered in the Saga Pattern article. This page is about the choice that comes next and is much harder to reverse: who decides what happens after each step?

In choreography, nobody decides. Each service listens for events, does its local work and emits a new event; the saga is the emergent chain. In orchestration, one component, the orchestrator, holds the saga's state and sends commands to each participant in turn. Both deliver the same business outcome on the happy path. They differ in coupling, in what it costs to change the flow, in how you notice a stuck saga, and in who gets paged. This article builds one saga both ways, compares them on those axes with concrete numbers, and finishes with a decision matrix, hybrid designs and a migration path.

One saga, two ways

The running example is order checkout with three participants: Inventory reserves stock, Payment authorizes the card, Shipping creates a shipment. Payment is the pivot; once the card is authorized and the shipment created, the order goes forward. If Payment fails, the stock reservation must be released. If Shipping fails, the authorization must be voided and the stock released.

Choreography: events, no ownerOrderOrderPlacedInventoryStockReservedPaymentPaymentAuthorizedShippingShipmentCreatedOrder listensflow lives in four sets of subscriptionsOrchestration: commands, one ownerOrder sagastate machine + timersInventoryPaymentShippingReserveAuthorizeCreateflow lives in one file; replies return to the sagaSame steps, same compensations; different place where the sequence is written down
Figure 1. The same checkout saga. Left: each service reacts to the previous service's event. Right: an orchestrator sends commands and receives replies.

Both versions need the same foundations: each local transaction publishes its message atomically with its state change, usually through a transactional outbox; every handler is idempotent because delivery is at least once; and every message carries the saga's correlation ID, here the order ID. What differs is where the sequence is written down.

Choreography in code

In the choreographed version, each service knows which events it reacts to and which events it emits. The sequence exists only as the union of these subscriptions.

# inventory service
@on("OrderPlaced")
def reserve(evt, tx):
    if tx.seen(evt.id): return                     # idempotency via processed-message table
    ok = tx.reserve_stock(evt.order_id, evt.lines)
    tx.outbox.emit("StockReserved" if ok else "StockRejected", order_id=evt.order_id)

@on("PaymentDeclined")
@on("ShipmentFailed")
def release(evt, tx):                              # compensation, triggered by others' failures
    if tx.seen(evt.id): return
    tx.release_stock(evt.order_id)
    tx.outbox.emit("StockReleased", order_id=evt.order_id)

# payment service
@on("StockReserved")
def authorize(evt, tx): ...                        # emits PaymentAuthorized or PaymentDeclined

@on("ShipmentFailed")
def void(evt, tx): ...                             # emits PaymentVoided

# order service
@on("ShipmentCreated")   -> mark CONFIRMED
@on("StockRejected")     -> mark REJECTED
@on("PaymentDeclined")   -> mark PAYMENT_FAILED

Notice what Inventory has to know: that a declined payment and a failed shipment both mean it must release stock. Its compensation logic is coupled to the failure events of two other services. Notice also that no component knows the order is stuck if StockReserved is published and Payment never reacts. The order service only sees terminal events, so the saga simply stops.

Orchestration in code

In the orchestrated version, participants expose commands and reply with outcomes; they do not know what comes before or after them. The orchestrator persists a state machine per order. Building the orchestrator itself is covered separately; the shape is this:

STEPS = [
    # (command,           participant,  compensation)
    ("ReserveStock",      "inventory",  "ReleaseStock"),
    ("AuthorizePayment",  "payment",    "VoidAuthorization"),
    ("CreateShipment",    "shipping",   None),
]

def on_reply(saga, reply, tx):                 # saga row loaded FOR UPDATE
    if reply.id in saga.handled: return
    saga.handled.add(reply.id)
    if reply.ok and saga.state == "RUNNING":
        saga.step += 1
        if saga.step == len(STEPS):
            saga.state = "DONE"
        else:
            send(STEPS[saga.step], saga, tx)    # via outbox; sets deadline timer
    else:
        saga.state = "COMPENSATING"
        for cmd, who, comp in reversed(STEPS[:saga.step]):
            if comp: tx.outbox.command(who, comp, order_id=saga.id)

def on_timer(saga, tx):                         # deadline passed with no reply
    if saga.retries < 3: resend_current(saga, tx); saga.retries += 1
    else: on_reply(saga, Reply(id=f"timeout-{saga.step}", ok=False), tx)

Inventory now has two commands, reserve and release, and no idea why release is called. The whole flow, including the compensation order and the timeout policy, is in one place you can read, test and version.

What it costs to change the flow

The clearest way to compare the two is to change the flow. Suppose the business adds a fraud check between stock reservation and payment, and a fraud rejection must release stock.

ChangeChoreographyOrchestration
New step: FraudCheck after stock reservedFraud subscribes to StockReserved; Payment must switch from StockReserved to FraudClearedAdd one row to STEPS; Fraud exposes a command
Compensation for fraud rejectionInventory adds a handler for FraudRejected; Order adds a terminal stateCovered by the existing reverse loop
Services deployed4 (Fraud, Payment, Inventory, Order)2 (Fraud, orchestrator)
Deploy ordering hazardIf Payment switches before Fraud ships, payments stop; if after, some orders skip fraudOrchestrator version pin per saga; in-flight sagas keep old steps
Tests that prove the new flowCross-service contract tests plus an end-to-end runUnit tests on the state machine plus one contract per participant

The deploy-ordering row is the one teams underestimate. In choreography the flow is a distributed configuration, so changing it is a coordinated multi-service release. In orchestration it is a code change in one service, and in-flight sagas can be pinned to the version of the step list they started with.

Stuck sagas: who notices?

Every saga eventually meets a participant that is down, slow or buggy. The question is who notices.

Orchestration answers by construction. Each step sets a deadline; the orchestrator retries, then compensates, and a query on its table (state = 'RUNNING' AND updated_at < now() - interval '10 minutes') lists every stuck saga. The cost is that the orchestrator is now a critical dependency. It must be durable, horizontally scaled by partitioning sagas by ID, and highly available, or every flow stops when it does.

Choreography has no owner of time, so you must add one. The usual answer is a saga monitor: a read-only consumer that subscribes to every event in the flow, keeps the last event per correlation ID, and alerts when an order has been in a non-terminal state for too long.

TERMINAL = {"ShipmentCreated", "StockRejected", "PaymentDeclined", "StockReleased"}
EXPECTED_NEXT_WITHIN = {"OrderPlaced": 60, "StockReserved": 120, "PaymentAuthorized": 300}

@on_any(topics=["orders", "inventory", "payments", "shipping"])
def track(evt, db):
    db.upsert("saga_progress", key=evt.order_id,
              last=evt.type, at=evt.time, terminal=evt.type in TERMINAL)

def sweep(db, now):                                  # every minute
    for row in db.query("SELECT * FROM saga_progress WHERE NOT terminal"):
        limit = EXPECTED_NEXT_WITHIN.get(row.last)
        if limit and (now - row.at).total_seconds() > limit:
            alert(f"order {row.key} stuck after {row.last}")

Look at what that monitor contains: the list of events, the terminal states and the expected timing of each transition. It is the flow definition again, written a second time and kept in sync by hand. Teams that run choreography at scale end up with this component, and it is often the first step toward orchestration.

Choreography also has a failure mode orchestration cannot have: event cycles. If Payment emits PaymentRetried and Inventory reacts by re-reserving, which Payment treats as a new trigger, two well-meaning handlers can loop. Draw the event graph and check it for cycles whenever a subscription is added.

Observability: where is order 8812?

Answering 'where is order 8812?' should take one query. With an orchestrator it does: the saga row holds the state, the current step, the retry count and the history of replies. With choreography the answer is spread across four services' logs. Make it reconstructable from day one: put the correlation ID and a causation ID (the ID of the event that triggered this one) in every message header, propagate trace context so one trace covers the whole saga, and build the monitor above even if you do not alert from it yet. Without these, a choreographed saga is debuggable only by the people who wrote it.

The causation ID is what turns a pile of events into a tree. Given every event for order 8812, sorting by time tells you the order in which things were recorded, which can differ from the order in which they caused each other when clocks drift or a retry lands late. Following causation links tells you that StockReleased was caused by PaymentDeclined, which was caused by StockReserved. Support tooling can render that tree directly, and it is the same in both styles, so it is worth building even if you orchestrate.

Testing differs in the same way. An orchestrator's state machine is a pure function from (state, reply) to (new state, commands), so you can unit-test every failure path, including a timeout at each step and a duplicate reply, in milliseconds. A choreographed flow has no single function to test; its behaviour only exists when all participants run together. Teams cover it with consumer-driven contract tests for each event plus a small number of end-to-end runs in a shared environment, and those runs are slower, flakier and usually the first thing skipped under deadline pressure. Count that cost when you choose.

Decision matrix

FactorFavours choreographyFavours orchestration
Steps in the flow2 or 34 or more, or growing
Teams owning participantsOne team, or loosely related teamsSeveral teams that release independently
Flow change frequencyRareFrequent business-rule changes
Need to see per-instance stateLow; aggregate metrics sufficeHigh; support staff ask about individual orders
Timeouts and deadlinesFew or noneMany, with business meaning (hold expires in 15 minutes)
Other consumers of the same eventsMany, and they should not be coupled to the flowFew
Tolerance for a central dependencyLowAcceptable if it is durable and partitioned

A rule of thumb that holds up: if you can describe the flow as 'when X happens, Y reacts' and nobody needs to ask where a single instance is, choreography is lighter. If the flow has an owner in the business, a deadline, or more than three steps, give it an owner in the code. Managed workflow engines, such as AWS Step Functions, are a way to get orchestration without building the durable state machine yourself.

Hybrids and migration

Real systems mix the two, and the clean seam is the bounded context. Inside one context, where one team owns the flow, orchestrate. Between contexts, publish domain events and let other contexts react. In the example, an Order saga orchestrates reserve, fraud, authorize and ship; when it finishes it publishes OrderConfirmed, and Loyalty, Analytics and Email react on their own without the saga knowing they exist. The saga has an owner; the rest of the company stays decoupled.

The orchestrator itself needs the same care as any stateful service. Store saga state in a database the orchestrator owns, partition it by saga ID so several instances can run at once without two of them advancing the same saga, and drive timers from a persisted deadline column rather than in-memory schedulers that vanish on restart. When the orchestrator is down, participants keep working and their replies queue up; when it comes back, it resumes from stored state. That property, recoverable by replaying replies against persisted state, is what makes the central dependency acceptable.

Migrating an existing choreographed flow is safest in three steps. First, deploy the orchestrator in shadow mode: it consumes the existing events, advances its own state machine and records what it would have commanded, without sending anything. Compare its view with reality for a few weeks; disagreements reveal undocumented paths. Second, move one transition at a time: have the orchestrator send the command and switch that participant from the event subscription to the command, behind a flag. Third, delete the old subscriptions and keep the domain events for outside consumers. Distributed Transactions in Practice covers when a saga is the wrong tool altogether.

What to do next

  1. Write down your saga as a step table: command, participant, compensation, deadline. If it does not fit on one screen, lean toward orchestration.
  2. Put a correlation ID and a causation ID in every saga message today, whichever style you use.
  3. If you choreograph, build the saga monitor and draw the event graph; check it for cycles on every new subscription.
  4. If you orchestrate, partition saga state by ID, store per-step deadlines, and alert on sagas past their deadline.
  5. Make every participant handler idempotent and publish through an outbox.
  6. Use the decision matrix with your actual step count, team count and change rate, not a preference.
  7. For an existing choreographed flow that hurts, run an orchestrator in shadow mode before moving any transition.
Key takeaway: Choreography and orchestration run the same steps and compensations; they differ in where the sequence is written down. Choreography keeps services decoupled for short, stable flows but spreads the flow across subscriptions, so changes need coordinated releases and stuck sagas need a monitor you must build. Orchestration puts the flow, deadlines and per-instance state in one durable component at the cost of a central dependency. Orchestrate inside a bounded context and publish events between contexts.