Auto Scaling and Load Balancing

Add and remove EC2 capacity automatically and spread traffic across it with Elastic Load Balancing.

What is it?

Elasticity means capacity follows demand. Two services work together to deliver it: Amazon EC2 Auto Scaling changes the number of instances, and Elastic Load Balancing (ELB) spreads incoming requests across whichever instances exist.

An Auto Scaling group has three key numbers:

  • Minimum: the fewest instances ever running.
  • Desired: how many you want right now.
  • Maximum: the ceiling, which protects your budget.

Scaling approaches:

  • Dynamic scaling: react to metrics, for example 'keep average CPU near 50%' (target tracking) or step/simple policies tied to alarms.
  • Predictive scaling: use machine learning on past patterns to scale ahead of a recurring peak.
  • Scheduled scaling: scale at known times, like a Monday-morning rush.

Load balancer types:

  • Application Load Balancer (ALB): layer 7 (HTTP/HTTPS). Routes by path, host, headers; good for web apps and microservices.
  • Network Load Balancer (NLB): layer 4 (TCP/UDP/TLS). Very high throughput, low latency, static IPs.
  • Gateway Load Balancer (GWLB): sends traffic through third-party virtual appliances such as firewalls and inspection tools.
  • Classic Load Balancer is the legacy option.

Load balancers also run health checks and stop sending traffic to unhealthy targets.

Explain like I'm 10

Imagine an airport check-in hall. The load balancer is the person at the entrance who points each traveler to the shortest open desk and stops pointing at a desk whose agent stepped away. Auto Scaling is the manager who opens more desks when the line grows and closes them when it empties - never fewer than two, never more than the hall can hold.

Examples

Auto Scaling group with target tracking (CLI)

aws autoscaling create-auto-scaling-group \
  --auto-scaling-group-name web-asg \
  --launch-template LaunchTemplateName=web-template,Version='$Latest' \
  --min-size 2 --desired-capacity 2 --max-size 6 \
  --vpc-zone-identifier "subnet-aaa111,subnet-bbb222" \
  --target-group-arns arn:aws:elasticloadbalancing:eu-west-1:111122223333:targetgroup/web/abc123 \
  --health-check-type ELB --health-check-grace-period 120

aws autoscaling put-scaling-policy \
  --auto-scaling-group-name web-asg \
  --policy-name cpu50 \
  --policy-type TargetTrackingScaling \
  --target-tracking-configuration '{
    "PredefinedMetricSpecification": {"PredefinedMetricType": "ASGAverageCPUUtilization"},
    "TargetValue": 50.0
  }'

Two subnets in different AZs give you multi-AZ capacity; the ELB health check replaces instances that stop answering.

Which load balancer?

Need                                           Choose
---------------------------------------------  -----
Route /api to one service, /img to another     ALB
HTTP/2, WebSockets, host-based routing         ALB
Millions of TCP connections, fixed IPs         NLB
Insert a firewall appliance into the path      GWLB

How it works

Metrics (such as CPU or request count per target) flow to CloudWatch. A scaling policy compares the metric with its target and changes the desired capacity within min and max. The group launches instances from a launch template or terminates the least useful ones, balancing across AZs.

The load balancer sits in front, registers new instances in a target group, health-checks them, and distributes requests. Because instances are interchangeable, the application should keep session data outside the server (for example in a cache or database).

         users
           |
   [ Load Balancer ]  <- health checks
      /      |      \
  +-----+ +-----+ +-----+
  | EC2 | | EC2 | | EC2 |   Auto Scaling group
  +-----+ +-----+ +-----+   min 2 / desired 3 / max 6
   AZ-a    AZ-b    AZ-a
       ^ CloudWatch metric drives desired count

Why does it exist?

Fixed capacity is wasteful when traffic is low and an outage when traffic spikes. Together these services keep cost proportional to demand and keep the application available when individual servers fail.

When to use it

Use them for any stateless web or API tier with variable load, and whenever you need self-healing: a failed instance is replaced automatically.

When not to use it

A single stateful server (like a lone database on EC2) does not benefit from simply adding more copies. Very small, constant workloads may not need scaling at all. Serverless options like Lambda scale without you managing groups.

Common mistakes

  • Setting min = max = 1 and believing that is 'highly available'.

  • Placing all instances in a single AZ.

  • Storing user sessions on local disk, so scaling in loses data.

  • Health-check grace period too short, so new instances are killed while booting.

  • Using an ALB when you need static IPs (NLB) or the reverse.

Practice exercises

  1. Easy:

    Explain min, desired, and max for a group set to 2/4/10 and what happens if an instance crashes.

  2. Medium:

    Choose dynamic, predictive, or scheduled scaling for: (a) a payroll app on the 1st and 15th, (b) a viral campaign, (c) a daily commuter app.

  3. Medium:

    Write ALB path rules to send /api/* to one target group and everything else to another.

  4. Hard:

    Design a stateless login flow so any instance can serve any user. Where does the session live?

Interview questions

Difference between EC2 Auto Scaling and ELB?

Auto Scaling changes how many instances exist; ELB distributes traffic across them and checks their health.

When use an NLB over an ALB?

When you need layer 4 handling, extreme performance, or static IP addresses; use an ALB for HTTP-aware routing.

What is predictive scaling?

Forecasting load from historical patterns and scaling capacity ahead of time.

Exam-style: Which service automatically adds EC2 instances when demand rises?

Amazon EC2 Auto Scaling.

Exam-style: Which load balancer routes traffic based on URL path?

Application Load Balancer.