Messaging: SQS, SNS and EventBridge
Decouple components with queues and publish/subscribe topics so one failure does not cascade.
What is it?
In a tightly coupled design, service A calls service B directly and waits. If B is slow or down, A suffers. In a loosely coupled design, A drops a message somewhere durable and moves on; B picks it up when it can.
- Amazon SQS (Simple Queue Service): a managed queue. Producers send messages, consumers poll and process them, then delete them. Messages are stored until processed. Standard queues offer very high throughput with at-least-once delivery and best-effort ordering; FIFO queues give ordering and exactly-once processing semantics with lower throughput. A dead-letter queue collects messages that keep failing.
- Amazon SNS (Simple Notification Service): publish/subscribe. A publisher sends one message to a topic, and SNS pushes copies to every subscriber (SQS queues, Lambda, HTTPS endpoints, email, SMS, mobile push).
- Amazon EventBridge: an event bus that routes events from AWS services, SaaS apps, and your code to targets based on rules, with schedules and filtering.
Combining SNS and SQS gives fan-out: one event, many independent queues, each processed at its own pace.
Explain like I'm 10
A queue is a ticket spindle in a repair workshop: customers pin job slips, and any mechanic grabs the next slip when free; slips wait safely if everyone is busy. A topic is a notice-board announcement: one note, and every team that signed up for those announcements gets its own copy.
Examples
Queue basics with the CLI
QUEUE_URL=$(aws sqs create-queue --queue-name orders --query QueueUrl --output text)
aws sqs send-message --queue-url "$QUEUE_URL" \
--message-body '{"orderId":"A-1001","total":42.5}'
aws sqs receive-message --queue-url "$QUEUE_URL" \
--wait-time-seconds 10 --max-number-of-messages 1
# After processing, delete using the ReceiptHandle from the receive call:
# aws sqs delete-message --queue-url "$QUEUE_URL" --receipt-handle "<handle>"A received message becomes invisible for the visibility timeout; if you do not delete it, it reappears for another try.
Fan-out design (CloudFormation sketch)
Resources:
OrderEvents:
Type: AWS::SNS::Topic
BillingQueue:
Type: AWS::SQS::Queue
ShippingQueue:
Type: AWS::SQS::Queue
BillingSub:
Type: AWS::SNS::Subscription
Properties:
TopicArn: !Ref OrderEvents
Protocol: sqs
Endpoint: !GetAtt BillingQueue.Arn
ShippingSub:
Type: AWS::SNS::Subscription
Properties:
TopicArn: !Ref OrderEvents
Protocol: sqs
Endpoint: !GetAtt ShippingQueue.Arn
# A real template also needs a queue policy allowing SNS to send messages.One 'order placed' publish reaches billing and shipping independently.
How it works
SQS stores each message redundantly across AZs. A consumer receives it, the message turns invisible for a visibility timeout, and the consumer deletes it when done. If the consumer crashes, the timeout expires and the message is retried. After a configured number of failed receives, it moves to the dead-letter queue.
SNS delivers each published message to all subscriptions, with retries and optional filter policies so subscribers only get what they care about. EventBridge evaluates rules against event patterns and forwards matches to targets.
Producer --> [ SNS topic ] --+--> [ SQS billing ] --> Billing workers
+--> [ SQS shipping ] --> Shipping workers
+--> Lambda (audit)
Failed 3x ... --> [ Dead-letter queue ]Why does it exist?
Direct calls tie components' availability and speed together. Messaging absorbs bursts, hides slow consumers, enables retries, and lets teams change one component without breaking others.
When to use it
Use SQS to buffer work (image processing, order handling). Use SNS for notifications and fan-out. Use EventBridge for event-driven integrations across services and SaaS, rule-based routing, and scheduled events.
When not to use it
When the caller truly needs an immediate answer (a login check), use a synchronous call. For high-volume streaming analytics with replay, consider a streaming service such as Kinesis instead of a queue.
Common mistakes
Assuming Standard SQS never delivers a message twice - consumers must be idempotent.
Visibility timeout shorter than the processing time, causing duplicate work.
No dead-letter queue, so poison messages loop forever.
Using SNS when you need buffering for a slow consumer (add an SQS queue behind it).
Forgetting the access policy that lets SNS write into an SQS queue.
Practice exercises
- Easy:
Explain in your own words the difference between a queue and a pub/sub topic.
- Medium:
Create a queue and a dead-letter queue with a max receive count of 3. Send a message and fail to delete it to watch it move.
- Medium:
Design a fan-out for 'user signed up' that sends a welcome email, creates a CRM record, and logs analytics.
- Hard:
Make a consumer idempotent against duplicate delivery. What key would you store, and where?
Interview questions
SQS vs SNS?
SQS is a pull-based queue where one consumer group processes each message; SNS is push-based pub/sub that delivers copies to many subscribers.
What is a dead-letter queue?
A queue that receives messages that could not be processed successfully after a set number of attempts, so they can be inspected.
Standard vs FIFO queue?
Standard: highest throughput, at-least-once, best-effort order. FIFO: strict ordering within a group and deduplication, with lower throughput.
Exam-style: Which service decouples application components by storing messages until they are processed?
Amazon SQS.
Exam-style: Which service sends one message to many subscribers?
Amazon SNS.