Monitoring and Auditing
CloudWatch for performance, CloudTrail for who-did-what, Config for resource state, and Trusted Advisor for advice.
What is it?
Four services answer four different questions:
- Amazon CloudWatch - How is it performing? Collects metrics (CPU, request counts), logs, and events; lets you build dashboards and alarms that notify or trigger actions such as Auto Scaling.
- AWS CloudTrail - Who did what, when? Records API calls and console actions (management events, optionally data events) for auditing and investigation. CloudTrail Insights flags unusual activity patterns.
- AWS Config - What does my configuration look like and did it change? Records resource configurations over time and evaluates them against rules (for example 'all volumes must be encrypted').
- AWS Trusted Advisor - What could I improve? Checks your account against best practices in categories: cost optimization, performance, security, fault tolerance, service limits, and operational excellence. The number of available checks depends on your support plan.
Related services: AWS Health (notices about AWS events that affect you), X-Ray (tracing requests across services), and Systems Manager (operational management of fleets).
Explain like I'm 10
CloudWatch is the dashboard of a car: speed, fuel, warning lights. CloudTrail is the car's black-box recorder: every door opened and button pressed, with the driver's identity. Config is the maintenance log showing how the car was modified over time. Trusted Advisor is the mechanic who looks over the car and says 'your tires are worn'.
Examples
A CPU alarm
aws cloudwatch put-metric-alarm \
--alarm-name web-high-cpu \
--namespace AWS/EC2 --metric-name CPUUtilization \
--dimensions Name=AutoScalingGroupName,Value=web-asg \
--statistic Average --period 300 --evaluation-periods 2 \
--threshold 80 --comparison-operator GreaterThanThreshold \
--alarm-actions arn:aws:sns:eu-west-1:111122223333:ops-alertsIf average CPU stays above 80% for two 5-minute periods, notify the SNS topic.
Who deleted that bucket? (CloudTrail)
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=DeleteBucket \
--max-results 5 \
--query "Events[].[EventTime,Username,EventName]" --output tableLookup covers recent management events; for long-term analysis send the trail to S3.
A Config rule (CloudFormation)
Resources:
EncryptedVolumesRule:
Type: AWS::Config::ConfigRule
Properties:
ConfigRuleName: encrypted-volumes
Source:
Owner: AWS
SourceIdentifier: ENCRYPTED_VOLUMES
Scope:
ComplianceResourceTypes: ["AWS::EC2::Volume"]How it works
Services publish metrics to CloudWatch automatically (basic metrics at regular intervals; detailed monitoring is more frequent). Agents can add memory, disk, and application logs. Alarms evaluate metrics against thresholds and fire actions. CloudTrail delivers event logs to S3 and optionally CloudWatch Logs, and records identity, source IP, time, and request parameters. Config snapshots resource state and flags noncompliance.
Resources --metrics/logs--> CloudWatch --alarm--> SNS / Auto Scaling
Every API call ----------> CloudTrail --> S3 / audit
Resource config changes -> Config --rules--> compliant / non-compliant
Account best practices -> Trusted Advisor --> recommendationsWhy does it exist?
You cannot fix, secure, or optimise what you cannot see. Monitoring catches problems early, auditing answers 'who did that', and configuration tracking supports compliance and incident response.
When to use it
Turn on CloudTrail in every account, set alarms on key health metrics, use Config for compliance rules, and review Trusted Advisor regularly for cost and security wins.
When not to use it
CloudWatch is not an audit log of user actions (use CloudTrail), and CloudTrail is not a performance monitor (use CloudWatch). Do not treat Trusted Advisor as a complete security audit.
Common mistakes
Mixing up CloudWatch and CloudTrail.
Creating alarms nobody receives or acts on.
Not retaining CloudTrail logs long enough for investigations.
Alerting on too many noisy metrics until the team ignores them.
Assuming all Trusted Advisor checks are available on every support plan.
Practice exercises
- Easy:
Which service answers each: 'Why is the app slow?', 'Who changed the security group?', 'Is every volume encrypted?'
- Medium:
Create an alarm on a metric and an SNS email subscription. Trigger it and confirm the notification.
- Medium:
Use the cron builder tool to design a schedule for a nightly EventBridge rule that runs a cleanup job.
- Hard:
Outline an incident investigation using CloudWatch, CloudTrail and Config for 'an S3 bucket became public overnight'.
Interview questions
CloudWatch vs CloudTrail?
CloudWatch monitors performance and operational data (metrics, logs, alarms); CloudTrail records API activity for governance and audit.
What is AWS Config used for?
Recording resource configuration history and evaluating it against compliance rules.
What are Trusted Advisor categories?
Cost optimization, performance, security, fault tolerance, service limits, and operational excellence.
Exam-style: Which service logs API calls made in an AWS account?
AWS CloudTrail.
Exam-style: Which service can alarm when CPU utilisation exceeds a threshold?
Amazon CloudWatch.