A saga replaces one distributed transaction with a sequence of local transactions, each with a compensating action that undoes it if a later step fails. The mechanics of compensation, the pivot step and isolation are covered in the Saga Pattern article. This page is about the choice that comes next and is much harder to reverse: who decides what happens after each step?
In choreography, nobody decides. Each service listens for events, does its local work and emits a new event; the saga is the emergent chain. In orchestration, one component, the orchestrator, holds the saga's state and sends commands to each participant in turn. Both deliver the same business outcome on the happy path. They differ in coupling, in what it costs to change the flow, in how you notice a stuck saga, and in who gets paged. This article builds one saga both ways, compares them on those axes with concrete numbers, and finishes with a decision matrix, hybrid designs and a migration path.
One saga, two ways
The running example is order checkout with three participants: Inventory reserves stock, Payment authorizes the card, Shipping creates a shipment. Payment is the pivot; once the card is authorized and the shipment created, the order goes forward. If Payment fails, the stock reservation must be released. If Shipping fails, the authorization must be voided and the stock released.
Both versions need the same foundations: each local transaction publishes its message atomically with its state change, usually through a transactional outbox; every handler is idempotent because delivery is at least once; and every message carries the saga's correlation ID, here the order ID. What differs is where the sequence is written down.
Choreography in code
In the choreographed version, each service knows which events it reacts to and which events it emits. The sequence exists only as the union of these subscriptions.
# inventory service
@on("OrderPlaced")
def reserve(evt, tx):
if tx.seen(evt.id): return # idempotency via processed-message table
ok = tx.reserve_stock(evt.order_id, evt.lines)
tx.outbox.emit("StockReserved" if ok else "StockRejected", order_id=evt.order_id)
@on("PaymentDeclined")
@on("ShipmentFailed")
def release(evt, tx): # compensation, triggered by others' failures
if tx.seen(evt.id): return
tx.release_stock(evt.order_id)
tx.outbox.emit("StockReleased", order_id=evt.order_id)
# payment service
@on("StockReserved")
def authorize(evt, tx): ... # emits PaymentAuthorized or PaymentDeclined
@on("ShipmentFailed")
def void(evt, tx): ... # emits PaymentVoided
# order service
@on("ShipmentCreated") -> mark CONFIRMED
@on("StockRejected") -> mark REJECTED
@on("PaymentDeclined") -> mark PAYMENT_FAILEDNotice what Inventory has to know: that a declined payment and a failed shipment both mean it must release stock. Its compensation logic is coupled to the failure events of two other services. Notice also that no component knows the order is stuck if StockReserved is published and Payment never reacts. The order service only sees terminal events, so the saga simply stops.
Orchestration in code
In the orchestrated version, participants expose commands and reply with outcomes; they do not know what comes before or after them. The orchestrator persists a state machine per order. Building the orchestrator itself is covered separately; the shape is this:
STEPS = [
# (command, participant, compensation)
("ReserveStock", "inventory", "ReleaseStock"),
("AuthorizePayment", "payment", "VoidAuthorization"),
("CreateShipment", "shipping", None),
]
def on_reply(saga, reply, tx): # saga row loaded FOR UPDATE
if reply.id in saga.handled: return
saga.handled.add(reply.id)
if reply.ok and saga.state == "RUNNING":
saga.step += 1
if saga.step == len(STEPS):
saga.state = "DONE"
else:
send(STEPS[saga.step], saga, tx) # via outbox; sets deadline timer
else:
saga.state = "COMPENSATING"
for cmd, who, comp in reversed(STEPS[:saga.step]):
if comp: tx.outbox.command(who, comp, order_id=saga.id)
def on_timer(saga, tx): # deadline passed with no reply
if saga.retries < 3: resend_current(saga, tx); saga.retries += 1
else: on_reply(saga, Reply(id=f"timeout-{saga.step}", ok=False), tx)Inventory now has two commands, reserve and release, and no idea why release is called. The whole flow, including the compensation order and the timeout policy, is in one place you can read, test and version.
What it costs to change the flow
The clearest way to compare the two is to change the flow. Suppose the business adds a fraud check between stock reservation and payment, and a fraud rejection must release stock.
| Change | Choreography | Orchestration |
|---|---|---|
| New step: FraudCheck after stock reserved | Fraud subscribes to StockReserved; Payment must switch from StockReserved to FraudCleared | Add one row to STEPS; Fraud exposes a command |
| Compensation for fraud rejection | Inventory adds a handler for FraudRejected; Order adds a terminal state | Covered by the existing reverse loop |
| Services deployed | 4 (Fraud, Payment, Inventory, Order) | 2 (Fraud, orchestrator) |
| Deploy ordering hazard | If Payment switches before Fraud ships, payments stop; if after, some orders skip fraud | Orchestrator version pin per saga; in-flight sagas keep old steps |
| Tests that prove the new flow | Cross-service contract tests plus an end-to-end run | Unit tests on the state machine plus one contract per participant |
The deploy-ordering row is the one teams underestimate. In choreography the flow is a distributed configuration, so changing it is a coordinated multi-service release. In orchestration it is a code change in one service, and in-flight sagas can be pinned to the version of the step list they started with.
Stuck sagas: who notices?
Every saga eventually meets a participant that is down, slow or buggy. The question is who notices.
Orchestration answers by construction. Each step sets a deadline; the orchestrator retries, then compensates, and a query on its table (state = 'RUNNING' AND updated_at < now() - interval '10 minutes') lists every stuck saga. The cost is that the orchestrator is now a critical dependency. It must be durable, horizontally scaled by partitioning sagas by ID, and highly available, or every flow stops when it does.
Choreography has no owner of time, so you must add one. The usual answer is a saga monitor: a read-only consumer that subscribes to every event in the flow, keeps the last event per correlation ID, and alerts when an order has been in a non-terminal state for too long.
TERMINAL = {"ShipmentCreated", "StockRejected", "PaymentDeclined", "StockReleased"}
EXPECTED_NEXT_WITHIN = {"OrderPlaced": 60, "StockReserved": 120, "PaymentAuthorized": 300}
@on_any(topics=["orders", "inventory", "payments", "shipping"])
def track(evt, db):
db.upsert("saga_progress", key=evt.order_id,
last=evt.type, at=evt.time, terminal=evt.type in TERMINAL)
def sweep(db, now): # every minute
for row in db.query("SELECT * FROM saga_progress WHERE NOT terminal"):
limit = EXPECTED_NEXT_WITHIN.get(row.last)
if limit and (now - row.at).total_seconds() > limit:
alert(f"order {row.key} stuck after {row.last}")Look at what that monitor contains: the list of events, the terminal states and the expected timing of each transition. It is the flow definition again, written a second time and kept in sync by hand. Teams that run choreography at scale end up with this component, and it is often the first step toward orchestration.
Choreography also has a failure mode orchestration cannot have: event cycles. If Payment emits PaymentRetried and Inventory reacts by re-reserving, which Payment treats as a new trigger, two well-meaning handlers can loop. Draw the event graph and check it for cycles whenever a subscription is added.
Observability: where is order 8812?
Answering 'where is order 8812?' should take one query. With an orchestrator it does: the saga row holds the state, the current step, the retry count and the history of replies. With choreography the answer is spread across four services' logs. Make it reconstructable from day one: put the correlation ID and a causation ID (the ID of the event that triggered this one) in every message header, propagate trace context so one trace covers the whole saga, and build the monitor above even if you do not alert from it yet. Without these, a choreographed saga is debuggable only by the people who wrote it.
The causation ID is what turns a pile of events into a tree. Given every event for order 8812, sorting by time tells you the order in which things were recorded, which can differ from the order in which they caused each other when clocks drift or a retry lands late. Following causation links tells you that StockReleased was caused by PaymentDeclined, which was caused by StockReserved. Support tooling can render that tree directly, and it is the same in both styles, so it is worth building even if you orchestrate.
Testing differs in the same way. An orchestrator's state machine is a pure function from (state, reply) to (new state, commands), so you can unit-test every failure path, including a timeout at each step and a duplicate reply, in milliseconds. A choreographed flow has no single function to test; its behaviour only exists when all participants run together. Teams cover it with consumer-driven contract tests for each event plus a small number of end-to-end runs in a shared environment, and those runs are slower, flakier and usually the first thing skipped under deadline pressure. Count that cost when you choose.
Decision matrix
| Factor | Favours choreography | Favours orchestration |
|---|---|---|
| Steps in the flow | 2 or 3 | 4 or more, or growing |
| Teams owning participants | One team, or loosely related teams | Several teams that release independently |
| Flow change frequency | Rare | Frequent business-rule changes |
| Need to see per-instance state | Low; aggregate metrics suffice | High; support staff ask about individual orders |
| Timeouts and deadlines | Few or none | Many, with business meaning (hold expires in 15 minutes) |
| Other consumers of the same events | Many, and they should not be coupled to the flow | Few |
| Tolerance for a central dependency | Low | Acceptable if it is durable and partitioned |
A rule of thumb that holds up: if you can describe the flow as 'when X happens, Y reacts' and nobody needs to ask where a single instance is, choreography is lighter. If the flow has an owner in the business, a deadline, or more than three steps, give it an owner in the code. Managed workflow engines, such as AWS Step Functions, are a way to get orchestration without building the durable state machine yourself.
Hybrids and migration
Real systems mix the two, and the clean seam is the bounded context. Inside one context, where one team owns the flow, orchestrate. Between contexts, publish domain events and let other contexts react. In the example, an Order saga orchestrates reserve, fraud, authorize and ship; when it finishes it publishes OrderConfirmed, and Loyalty, Analytics and Email react on their own without the saga knowing they exist. The saga has an owner; the rest of the company stays decoupled.
The orchestrator itself needs the same care as any stateful service. Store saga state in a database the orchestrator owns, partition it by saga ID so several instances can run at once without two of them advancing the same saga, and drive timers from a persisted deadline column rather than in-memory schedulers that vanish on restart. When the orchestrator is down, participants keep working and their replies queue up; when it comes back, it resumes from stored state. That property, recoverable by replaying replies against persisted state, is what makes the central dependency acceptable.
Migrating an existing choreographed flow is safest in three steps. First, deploy the orchestrator in shadow mode: it consumes the existing events, advances its own state machine and records what it would have commanded, without sending anything. Compare its view with reality for a few weeks; disagreements reveal undocumented paths. Second, move one transition at a time: have the orchestrator send the command and switch that participant from the event subscription to the command, behind a flag. Third, delete the old subscriptions and keep the domain events for outside consumers. Distributed Transactions in Practice covers when a saga is the wrong tool altogether.
What to do next
- Write down your saga as a step table: command, participant, compensation, deadline. If it does not fit on one screen, lean toward orchestration.
- Put a correlation ID and a causation ID in every saga message today, whichever style you use.
- If you choreograph, build the saga monitor and draw the event graph; check it for cycles on every new subscription.
- If you orchestrate, partition saga state by ID, store per-step deadlines, and alert on sagas past their deadline.
- Make every participant handler idempotent and publish through an outbox.
- Use the decision matrix with your actual step count, team count and change rate, not a preference.
- For an existing choreographed flow that hurts, run an orchestrator in shadow mode before moving any transition.