Google Cloud Pub/Sub looks simple from the outside: you publish bytes to a topic, and every subscription on that topic gets its own copy. The behaviour that decides whether a pipeline built on it is correct sits underneath that simple interface. It decides when a publish counts as safe, what a lease on a message means, why a slow consumer sees duplicates, what an ordering key costs you, and what "exactly-once" actually promises. This article walks through the service the way it is built, a control plane that places clients and a data plane that moves messages, and then turns each piece into a configuration decision you can make deliberately.

By the end you should be able to size a subscriber from its processing time, choose between pull, push and export subscriptions, decide whether you need ordering keys or exactly-once delivery, and build a dead-letter and replay path that does not quietly lose or reorder data.

Pub/Sub: routers place clients, forwarders move messagesPublisherbatches, ordering keyPublishing forwardernearest data centerDurable storagereplicated, then ackSubscription Apull / StreamingPullSubscription Bpush to HTTPSpublishwritepublish acked only after the write is durableControl plane: routersassign clients to forwarders by network distance and loadSubscribing forwarderleases, ack deadlines, redeliverySubscriber clientflow control, lease extensionmessagesack / modackDead-letter topicafter max delivery attemptsRetention and seekreplay by timestamp or snapshotOnce connected, clients talk to forwarders directly; routers are not on the message path.
Routers in the control plane attach clients to nearby forwarders; the data plane stores each message durably and tracks leases per subscription.

Two planes: routers place, forwarders move

Google's architectural overview splits Pub/Sub into two parts. The data plane moves messages between publishers and subscribers, and its servers are called forwarders. The control plane decides which forwarders each client talks to, and its servers are called routers. When a client connects, a router picks data centers by shortest network distance and then spreads load across the forwarders inside them.

Two consequences matter for design. First, once a publisher or subscriber is attached to its forwarders, it does not need the routers again while those forwarders stay reachable, so control plane upgrades or trouble do not touch connected clients. Second, a topic is a global resource but the path a message takes is local: a publisher in one region writes through a nearby forwarder, and a subscriber in another region reads through its own nearby forwarder, which collects messages for that subscription from wherever they were stored. You get a global topic without running cross-region replication yourself, but latency and some guarantees, notably exactly-once, depend on where clients connect.

Two planes: routers place, forwarders move

Google's architectural overview splits Pub/Sub into two parts. The data plane moves messages between publishers and subscribers, and its servers are called forwarders. The control plane decides which forwarders each client talks to, and its servers are called routers. When a client connects, a router picks data centers by shortest network distance and then spreads load across the forwarders inside them.

Two consequences matter for design. First, once a publisher or subscriber is attached to its forwarders, it does not need the routers again while those forwarders stay reachable, so control plane upgrades or trouble do not touch connected clients. Second, a topic is a global resource but the path a message takes is local: a publisher in one region writes through a nearby forwarder, and a subscriber in another region reads through its own nearby forwarder, which collects messages for that subscription from wherever they were stored. You get a global topic without running cross-region replication yourself, but latency and some guarantees, notably exactly-once, depend on where clients connect.

The publish path: batching, durability and duplicates

A publish is acknowledged to the client only after the message has been written durably, and the response carries a server-assigned message ID. Everything before that moment is client-side: the library buffers messages and sends a batch when it reaches a count, a byte size or a delay threshold. The hard limits are 10 MB per message, 10 MB per publish request and 1,000 messages per request. Batching trades a few milliseconds of latency for far fewer requests, and on high-volume topics it is usually the single biggest throughput lever.

Publishing is at-least-once. If a publish times out, the client cannot know whether the write landed, so it retries, and a retry that duplicates a stored message produces a second message with a new message ID. Nothing later in the pipeline can recognise that duplicate by ID. If duplicates matter, put a business key such as an order ID and event version in an attribute at the source.

import json
from google.cloud import pubsub_v1
from google.cloud.pubsub_v1 import types

publisher = pubsub_v1.PublisherClient(
    batch_settings=types.BatchSettings(
        max_messages=500,        # send when 500 messages are buffered
        max_bytes=1_000_000,     # ... or 1 MB is buffered
        max_latency=0.02,        # ... or 20 ms have passed
    ),
    publisher_options=types.PublisherOptions(enable_message_ordering=True),
)
topic = publisher.topic_path("my-project", "orders")

def publish_order_event(event: dict) -> None:
    future = publisher.publish(
        topic,
        json.dumps(event).encode(),
        ordering_key=event["customer_id"],     # order is kept per customer only
        event_key=f'{event["order_id"]}:{event["version"]}',  # dedupe key for consumers
    )
    try:
        future.result(timeout=30)
    except Exception:
        # With ordering on, a failed publish pauses this key so later
        # messages cannot overtake it. Fix the cause, then resume.
        publisher.resume_publish(topic, event["customer_id"])
        raise

Subscriptions as independent cursors

A subscription is an independent cursor over the topic: each one receives every message published after it was created, keeps its own record of what is acknowledged and redelivers whatever is not. Adding a subscription never slows the others. Unacknowledged messages are kept for 7 days by default; with topic retention configured, messages can be kept for up to 31 days, which lets a new subscription seek back to data published before it existed.

Delivery typeHow messages moveUse it when
Pull / StreamingPullClient holds a long-lived gRPC stream; the library manages leases and acksWorkers on GKE or Compute Engine; high throughput; you need exactly-once
PushPub/Sub sends an HTTPS POST to your endpoint; a success status code is the ackCloud Run or other HTTP services that scale on request count
Export (BigQuery, Cloud Storage)Pub/Sub writes directly into the sinkLanding raw events with no transformation code to run

Push subscriptions control their own sending rate, backing off when your endpoint returns errors or slows down, and can attach an OIDC token so the endpoint can verify the caller. For per-request autoscaling see Cloud Run concurrency and autoscaling; if you want routing by event type across many Google sources, Eventarc sits on top of Pub/Sub.

Leases, ack deadlines and flow control

Delivery to a subscriber is a lease. When a message is handed out, the subscription starts its ack deadline (10 seconds by default, configurable up to 600). If the deadline passes without an ack, the message becomes eligible for redelivery, possibly to a different client. The client libraries extend leases automatically for messages they are still holding, by sending modify-ack-deadline requests, up to a configurable maximum lease duration.

Leases interact with flow control, the client-side limit on outstanding messages and bytes. If flow control lets a client take 10,000 messages and it can process 100 per second, the last messages wait 100 seconds in local memory while their leases are extended. Size the limit to what the worker can finish within a few seconds and let the rest stay on the server, where another worker can take them. The sizing rule is Little's law: messages in flight equals arrival rate times processing time.

from google.cloud import pubsub_v1
from google.cloud.pubsub_v1 import types

subscriber = pubsub_v1.SubscriberClient()
sub_path = subscriber.subscription_path("my-project", "orders-fulfilment")

def callback(message):
    key = message.attributes.get("event_key")
    if already_applied(key):               # idempotency on the business key,
        message.ack()                      # not on message.message_id
        return
    try:
        apply_in_transaction(key, message.data)   # write effect + key together
        message.ack()
    except TransientError:
        message.nack()                     # redeliver soon, counts as an attempt
    # any other exception: no ack, lease lapses, message is redelivered

flow = types.FlowControl(max_messages=200, max_bytes=50 * 1024 * 1024)
streaming_pull = subscriber.subscribe(sub_path, callback=callback, flow_control=flow)
streaming_pull.result()                    # block; cancel() on shutdown

Ordering keys and exactly-once, precisely

Ordering keys. With ordering enabled on the subscription and messages published with an ordering key in the same region, Pub/Sub delivers messages for that key in publish order. Messages for different keys stay fully parallel. The cost is that each key behaves like a single-lane queue: a subscriber cannot work on the next message for a key until the previous one is acked, and redelivery of one message brings the messages after it on that key back as well. Publisher throughput is limited to 1 MBps per ordering key. Choose keys with many distinct values, such as a customer or account, never a constant or a low-cardinality field such as region.

Exactly-once delivery. This is a subscription setting, and its promise is narrower than its name. When enabled, a message whose ack succeeded will not be redelivered, and the subscriber can learn whether each ack succeeded, which plain at-least-once subscriptions do not report. It applies only to pull subscriptions, including StreamingPull, and only when the subscribers connect to the service in the same region. The default ack deadline is 60 seconds rather than 10, and publish-to-subscribe latency is significantly higher. It does nothing about publisher retries, which create new message IDs, and nothing about side effects that happened before an ack failed. Treat it as a way to cut redelivery noise, and keep idempotent processing anyway.

# Python: confirm the ack really landed before treating the work as final
ack_future = message.ack_with_response()
try:
    ack_future.result()
except Exception:
    # The ack failed (for example the lease had already expired);
    # expect a redelivery and rely on idempotent processing.
    pass

Dead letters, retries and replay

A dead-letter policy moves a message to another topic after a maximum number of delivery attempts, which must be between 5 and 100 (default 5). Attempts are approximate, counted as nacks and deadline expiries. Two details catch teams out. First, the Pub/Sub service agent needs publisher rights on the dead-letter topic and subscriber rights on the source subscription; without them, forwarding silently fails and poison messages keep cycling. Second, a dead-letter topic with no subscription simply drops what is sent to it, so attach a subscription with long retention first.

Pair it with a retry policy using exponential backoff so a failing message does not hammer the consumer, and with seek: a snapshot taken before a risky deploy lets you rewind the subscription and reprocess, provided processing is idempotent.

gcloud pubsub topics create orders-dlq
gcloud pubsub subscriptions create orders-dlq-hold --topic=orders-dlq \
    --message-retention-duration=7d

gcloud pubsub subscriptions create orders-fulfilment --topic=orders \
    --ack-deadline=60 \
    --enable-message-ordering \
    --dead-letter-topic=orders-dlq --max-delivery-attempts=5 \
    --min-retry-delay=10s --max-retry-delay=600s

# before a risky consumer deploy
gcloud pubsub snapshots create pre-deploy --subscription=orders-fulfilment
# after a bad deploy: rewind and reprocess
gcloud pubsub subscriptions seek orders-fulfilment --snapshot=pre-deploy

Replaying dead letters by republishing them to the main topic gives them new IDs and puts them behind newer messages for the same key, so per-key order is not preserved across a replay. If ordering matters, the consumer must compare event versions and discard stale updates rather than assume arrival order.

Worked example: an order stream with a hot customer

An order service publishes 2,000 events per second, about 1 KB each, with the customer ID as ordering key. The fulfilment subscriber spends 40 ms per event, mostly a database write. By Little's law it needs 2,000 × 0.04 = 80 messages in flight. With four worker pods, flow control of 40 messages per pod leaves headroom without hoarding.

Now look at a single key. Because messages for one customer are processed one at a time, the maximum rate per key is 1 / 0.04 = 25 events per second. Ordinary customers send a few per minute, but one marketplace account sends 300 per second during a sale. Its backlog grows by 275 events every second no matter how many pods you add, its oldest-unacked age climbs, and the other customers look healthy. The fix is in the key design, not the subscriber: key by order ID instead of customer, since only events for the same order must be ordered, or shard the big account's key into a few sub-keys and reconcile by version.

Next, a schema change makes 0.1% of events fail parsing. With a dead-letter policy at 5 attempts, each bad event is tried five times with backoff and then parked. Because ordering is on, each affected customer's key is held back for that whole retry ladder, while every other key keeps flowing; 2 events per second land in the dead-letter hold subscription and an alert on its backlog fires. The engineer fixes the parser, and a small job republishes the parked events. The consumer's version check drops the few that are now stale. The pipeline never stalled for unaffected customers, and nothing was lost.

Failure modes

  • Duplicate side effects. Deduplicating on message ID misses publisher-retry duplicates. Fix: dedupe on a business key stored in the same transaction as the effect.
  • Hot ordering key. One key's backlog grows while total throughput looks fine. Fix: narrower keys; watch oldest unacked message age, not only backlog count.
  • Hoarding subscribers. Large flow-control limits strand messages in one pod while others idle, and expired leases cause redelivery storms. Fix: size flow control from Little's law.
  • Silent dead-lettering failure. Missing service-agent permissions or a dead-letter topic with no subscription. Fix: grant both roles and test with a deliberately bad message.
  • Push endpoint overload. A slow endpoint gets backed off, and backlog builds. Fix: return errors quickly, scale on concurrency, or switch to pull for heavy processing.
  • Poison message on an ordered key. Retries and backoff block every later message for that key until it is dead-lettered. Fix: fewer attempts and shorter backoff on ordered subscriptions.
  • Assuming exactly-once covers everything. Push subscriptions, cross-region subscribers and publisher retries are all outside it. Fix: idempotency stays mandatory.

Trade-offs

DecisionOption AOption B
OrderingOff: full parallelism, consumer handles order by versionOn: per-key order, per-key throughput ceiling
Delivery guaranteeAt-least-once: lowest latencyExactly-once: fewer redeliveries, pull only, higher latency
Delivery typePull: you control concurrency and leasesPush: no workers to run, Pub/Sub sets the pace
TransformationSubscriber code or DataflowExport subscription straight to BigQuery

If you are comparing services, the same ideas appear in SQS as visibility timeouts and FIFO message groups; see Amazon SQS architecture. The consumer-side discipline is covered in Idempotency in system design.

What to do next

  1. List every subscription with its delivery type, ack deadline, ordering and dead-letter settings.
  2. Measure processing time per message and set flow control from arrival rate times processing time.
  3. Put a business dedupe key in an attribute at publish time and make every consumer idempotent on it.
  4. Audit ordering keys: confirm high cardinality and compute the per-key ceiling of 1 / processing time.
  5. Add a dead-letter topic with a retained hold subscription, grant the service agent roles, and test it.
  6. Alert on oldest unacked message age and on dead-letter backlog, not only on message counts.
  7. Take a snapshot before risky consumer deploys and rehearse a seek-and-reprocess once.
Key takeaway: Pub/Sub publishes durably but at-least-once, delivers to each subscription through leases that lapse into redelivery, and orders only within a key, at a per-key throughput cost. Size flow control from Little's law, dedupe on business keys, keep ordering keys narrow, treat exactly-once as noise reduction rather than a correctness guarantee, and test the dead-letter and replay path before you need it.