Amazon Simple Notification Service (SNS) is a push-based publish/subscribe service. A producer publishes a message to a topic once, and SNS delivers a copy to every subscription on that topic: SQS queues, Lambda functions, HTTP endpoints, Firehose streams, email, SMS and mobile push. The producer does not know who the consumers are, and adding a consumer does not require touching the producer. That one property, decoupled fan-out, is why SNS sits in the middle of so many AWS architectures.
It is also easy to misuse. SNS standard topics do not keep messages for consumers that are not ready, they deliver at least once and without strict ordering, and a misconfigured filter or permission drops messages silently. This article explains the model from first principles, walks through publishing, filtering, retries and FIFO topics, builds an order-events fan-out with real code, and finishes with the failure modes and a checklist. Limits quoted here were checked against the AWS documentation in October 2026; check the current quotas page before relying on a number.
The model: topics, subscriptions and push delivery
A topic is a named channel identified by an ARN. A subscription binds a topic to one endpoint with a protocol, and carries its own settings: a filter policy, a delivery policy where the protocol allows one, raw message delivery and a redrive policy pointing at a dead-letter queue. When a message arrives, SNS evaluates every subscription independently and pushes to each one that matches. There is no consumer group and no offset: SNS is the one doing the work of delivery, unlike SQS or Kinesis, where consumers pull.
Two consequences follow. First, a standard topic is not a buffer. If an endpoint is down, SNS retries according to the delivery policy and then discards the message unless a DLQ is attached. The usual way to get durable, replayable consumption is to subscribe an SQS queue and let the consumer read from that. Second, delivery to each subscriber is independent, so one slow HTTP endpoint does not hold back the queue next to it.
The publish path and size limits
A producer calls Publish with a topic ARN, a message body and optional message attributes, or PublishBatch with up to ten entries. SNS stores the message redundantly before acknowledging, so a successful response means the topic has it; it does not mean any subscriber has received it. Treat a publish error or timeout as "unknown" and retry with the same business identifier, because the earlier attempt may have succeeded.
By default a topic accepts messages up to 256 KiB, counting the body and the attributes together. In September 2026 AWS added a MaximumMessageSize topic attribute that can be set anywhere from 1,024 to 1,048,576 bytes (1 MiB). A topic set above 256 KiB may only have Firehose, SQS and Lambda subscriptions, and at most 100 of them; HTTP, email, SMS and mobile push need the topic to stay at 256 KiB or below. For anything larger, the SNS Extended Client Libraries for Java and Python store the payload in S3 and publish a reference. In practice the better design is usually to publish a small event with an identifier and let consumers fetch what they need.
Subscriptions, protocols and the envelope
Subscribing an SQS queue or Lambda function in the same account is confirmed automatically once the resource policy allows it. HTTP, HTTPS and email subscriptions must be confirmed: SNS sends a SubscriptionConfirmation message containing a SubscribeURL, and nothing is delivered until the endpoint visits it. An HTTP endpoint that never handles that message type sits in "pending confirmation" forever, and every message to it is quietly skipped.
By default SNS wraps each message in a JSON envelope with fields such as Type, MessageId, TopicArn, Message, Timestamp, the attributes, and a signature. Consumers must parse the envelope and then parse Message. Enabling raw message delivery on an SQS, HTTP or Firehose subscription sends the body as-is, with attributes mapped to the target's own attributes where possible. Pick one per subscription and document it, because a consumer written for one shape fails on the other. HTTP endpoints should verify the signature using the certificate URL in the message, after checking that the URL belongs to an SNS domain.
Filter policies: routing without code
A filter policy is a JSON document on a subscription that decides which messages it receives. With the default scope, MessageAttributes, the policy matches message attributes. With FilterPolicyScope set to MessageBody, it matches fields inside a JSON body, including nested ones. Operators include exact values, prefix, suffix, anything-but, numeric ranges and exists. Filtering happens inside SNS, so filtered-out messages cost the subscriber nothing and never reach it.
{
"event_type": ["order_placed", "order_cancelled"],
"region": [{"prefix": "eu-"}],
"total": [{"numeric": [">=", 100]}],
"test": [{"exists": false}]
}Read this as AND across keys and OR within a key's list. The trap is what happens when the field is missing: a message without event_type does not match, and it is dropped without an error. If one producer forgets an attribute, every filtered subscriber stops getting its messages and only the NumberOfNotificationsFilteredOut metric shows it. Changes to filter policies are also eventually consistent, so a newly applied policy can take a short time before every message is evaluated against it. Never deploy a filter change and a producer change in the same step.
Delivery, retries and dead-letter queues
When delivery fails with a retryable error, SNS follows a delivery policy with four phases: immediate retries, a pre-backoff phase, a backoff phase and a post-backoff phase, with jitter applied. The policy depends on the protocol, and only HTTP and HTTPS can be customised.
| Endpoint type | Policy | What it means for you |
|---|---|---|
| SQS, Lambda, Firehose (AWS managed) | 3 immediate, 2 at 1 s, 10 backing off from 1 s to 20 s, then 100,000 at 20 s: 100,015 attempts over 23 days | Server-side errors retry for weeks, but a deleted queue or a broken permission is a client-side error that is not retried at all; alarm on failures |
| SMTP, SMS, mobile push | 0 immediate, 2 at 10 s, 10 backing off to 600 s, 38 at 600 s: 50 attempts over 6 hours | Treat these channels as best effort |
| HTTP and HTTPS | Customisable per topic or subscription; default is 3 retries 20 s apart; the whole policy is capped at 3,600 s | Set the policy to your server's real recovery time and set maxReceivesPerSecond to its capacity |
For HTTP endpoints, SNS retries 5XX and 429 responses. Any other error, such as a 400 or 403, is a permanent failure that is not retried. When the policy is exhausted, or the failure is permanent, the message is discarded unless the subscription has a redrive policy pointing at an SQS dead-letter queue. A subscription DLQ is the only record you will ever have of those messages, so attach one to every subscription that matters. The DLQ must sit in the same account and region as the subscription, its access policy must let SNS send to it, and for a FIFO topic it must be a FIFO queue.
FIFO topics: ordering and deduplication
A FIFO topic, whose name must end in .fifo, delivers messages in order within a message group and deduplicates them. Each publish carries a MessageGroupId, which defines the ordering scope, and a MessageDeduplicationId, or the topic enables content-based deduplication, which hashes the body. A duplicate published within the five-minute deduplication interval is accepted but not delivered again. A topic cannot be converted between standard and FIFO after it is created.
FIFO topics only deliver to SQS queues, standard or FIFO. HTTP, email, SMS and mobile push cannot subscribe, and to reach Lambda you subscribe a queue and let Lambda read from it. Ordering is end to end only if the subscribed queue is also FIFO. FIFO topics have lower throughput quotas than standard topics, and ordering is per group, so choose group keys with enough distinct values, such as an order or account ID, and never a constant. FIFO topics can also archive messages with a retention period and replay them to a subscription, which standard topics cannot do.
Worked example: order events fanned out to three services
An order service publishes one event per state change. Billing wants placed and cancelled orders, shipping wants paid orders, and an audit pipeline wants everything in S3. The producer publishes with attributes that the filters use and with an event ID that consumers use to drop duplicates, because standard topics deliver at least once.
import json, uuid, boto3
sns = boto3.client("sns")
TOPIC = "arn:aws:sns:eu-west-1:111122223333:orders"
def publish_order_event(order, event_type):
event_id = f"{order['id']}:{event_type}:{order['version']}" # stable across retries
body = {"event_id": event_id, "event_type": event_type,
"order_id": order["id"], "total": order["total"], "version": order["version"]}
return sns.publish(
TopicArn=TOPIC,
Message=json.dumps(body),
MessageAttributes={
"event_type": {"DataType": "String", "StringValue": event_type},
"region": {"DataType": "String", "StringValue": order["region"]},
},
)["MessageId"]The event ID is derived from the order and its version, not from uuid4, so a retried publish produces the same ID. Billing's queue subscription uses a filter on event_type, raw delivery and a DLQ. The queue's resource policy allows sqs:SendMessage from the service principal sns.amazonaws.com only when aws:SourceArn equals the topic ARN, so no other topic can write to it. The billing consumer records processed event IDs in a table with a conditional write and skips any it has already seen.
def handle(sqs_record, table):
event = json.loads(sqs_record["body"]) # raw delivery: no envelope
try:
table.put_item(Item={"event_id": event["event_id"]},
ConditionExpression="attribute_not_exists(event_id)")
except table.meta.client.exceptions.ConditionalCheckFailedException:
return # duplicate delivery, already handled
apply_billing(event) # must be safe if it fails after the putShipping subscribes a Lambda function with a filter for order_paid. Audit subscribes a Firehose stream with no filter. If shipping's function starts throwing, SNS's asynchronous invocation and Lambda's own retry settings apply, and billing and audit are unaffected. A useful test before go-live: publish one event of each type and check NumberOfNotificationsDelivered and NumberOfNotificationsFilteredOut per subscription against what you expected.
Security and encryption
Three policies decide whether a message flows. The topic's access policy controls who may publish and subscribe, and should name specific principals or use aws:SourceAccount and aws:SourceArn conditions when AWS services such as S3 or CloudWatch publish to it. The target's resource policy, on the queue, function or bucket, must allow SNS to deliver. If the topic uses server-side encryption with a customer managed KMS key, the key policy must let every publisher, including AWS service principals, use the key; publishes from services fail when it does not. If a subscribed queue is encrypted with a customer managed key, that key's policy must allow SNS too. A missing grant on the queue's key looks exactly like a message that was never sent.
What to watch in production
NumberOfNotificationsFailedper topic and subscription: the primary alarm. Any sustained non-zero value means messages are on their way to a DLQ or to nowhere.- Dead-letter queue depth via
ApproximateNumberOfMessagesVisibleon each DLQ, with an alarm at one message and a written redrive procedure. NumberOfNotificationsFilteredOutcompared withNumberOfMessagesPublished: a sudden jump means a producer changed an attribute.- Delivery status logging to CloudWatch Logs for HTTP, Lambda, SQS and Firehose subscriptions, sampled, to see response codes and dwell times.
- Publish errors and throttling on the producer side, since publish quotas are per account and region and differ by API.
Failure modes
- Silent filtering. A renamed attribute or a body field moved into a nested object stops delivery to every filtered subscriber with no errors.
- Pending confirmation. An HTTP subscriber that ignores the confirmation message never receives anything.
- Permanent HTTP errors. An endpoint that returns 400 during a deploy loses those messages immediately, because 4XX other than 429 is not retried.
- No DLQ. Retries end and the message is discarded, leaving only a metric.
- Duplicates and reordering on standard topics, which break consumers that assume exactly-once or in-order delivery.
- Envelope mismatch after someone toggles raw delivery, so consumers parse the wrong shape.
- KMS denials on an encrypted topic or queue that look like missing messages.
Trade-offs: SNS, EventBridge, SQS and Kinesis
Use SNS when you need simple, high-throughput fan-out to a known set of AWS targets or HTTP endpoints, or mobile and SMS notifications. Use EventBridge when routing rules are richer, events come from SaaS partners or AWS services, or you want schema discovery and archives on standard events, accepting lower default throughput and higher latency. Use SQS on its own for a single consumer group that needs buffering; use SNS in front of SQS when several groups each need their own buffered copy. Use Kinesis when consumers need ordered, replayable streams over hours or days. Lambda subscribers are covered in more depth in AWS Lambda.
What to do next
- List every topic and subscription and record the protocol, filter policy, raw delivery setting and whether a DLQ is attached.
- Attach an SQS dead-letter queue to every subscription that matters, with an alarm at one visible message and a redrive runbook.
- Add alarms on NumberOfNotificationsFailed and on a jump in NumberOfNotificationsFilteredOut.
- Make consumers idempotent on a stable event ID derived from business data, and test them with duplicated and reordered input.
- Restrict topic and queue policies with aws:SourceArn and aws:SourceAccount, and check KMS key policies for every publisher and subscriber.
- For HTTP subscribers, handle confirmation messages, verify signatures and set a delivery policy and maxReceivesPerSecond that match the server.
- Use a FIFO topic with SQS FIFO queues only where ordering or deduplication is a real requirement, and choose high-cardinality group IDs.