Load Balancing

Spreading incoming requests across multiple servers so no single one gets overwhelmed.

What is it?

A single server can only handle so many requests at once. Once traffic grows beyond that, you need more than one server — but then, how does a user's request know which server to go to? A load balancer sits in front of a group of servers and decides, for every incoming request, which one should handle it — spreading the work evenly so no single server gets overloaded while others sit idle.

Explain like I'm 10

A load balancer is like the host at a busy restaurant who seats guests across all the available tables, instead of letting one table get 20 parties while the rest sit empty.

Examples

Conceptual round-robin balancing

const servers = ["server-1", "server-2", "server-3"];
let next = 0;

function pickServer() {
  const server = servers[next];
  next = (next + 1) % servers.length; // cycle back to 0 after the last one
  return server;
}
// Requests get spread evenly: server-1, server-2, server-3, server-1, ...

This is a simplified version of 'round robin' — one of several strategies real load balancers use to distribute traffic.

How it works

Every incoming request first reaches the load balancer instead of a specific server directly. The load balancer picks a healthy server — using a strategy like round robin (cycling through servers in a fixed order, regardless of how busy each one currently is), least-connections (send each new request to whichever server currently has the fewest active requests still in progress), or based on server load — and forwards the request there. It also continuously checks server health, and stops sending traffic to any server that's down.

┌─────────────┐
Requests →│Load Balancer│
          └──────┬──────┘
        ┌────────┼────────┐
        ▼        ▼        ▼
    Server 1  Server 2  Server 3

Why does it exist?

Without load balancing, scaling up would mean building one enormous server — expensive, and still a single point of failure. Load balancing lets you add ordinary servers as traffic grows, and keeps the system running even if one individual server fails.

When to use it

Reach for a load balancer the moment a single server can no longer handle your traffic reliably, or when you want to survive one server failing without the whole application going down.

When not to use it

A single small server serving a low-traffic app doesn't need a load balancer yet — it adds infrastructure and cost for a problem you don't have. Add it when traffic or reliability requirements actually demand it, not preemptively for hypothetical scale.

Common mistakes

  • Treating the load balancer itself as unbreakable — it also needs redundancy, or it becomes a single point of failure.

  • Ignoring 'sticky sessions' — where the load balancer deliberately keeps routing a given user's later requests back to the same server it used before — when that user's data is only stored on that specific server.

  • Assuming load balancing alone solves scaling — the database or other shared resources behind it can still be the real bottleneck.

Practice exercises

  1. Easy:

    Explain, in your own words, why a single server isn't enough for a popular application.

  2. Medium:

    Describe the difference between 'round robin' and 'least connections' as load balancing strategies.

  3. Hard:

    Explain what could go wrong if a load balancer keeps sending traffic to a server that has crashed, and how health checks solve it.

Interview questions

What problem does a load balancer solve?

It distributes incoming requests across multiple servers so no single server is overwhelmed, and traffic keeps flowing even if one server fails.

What is a health check in load balancing?

A periodic check the load balancer performs on each server to confirm it's still responsive, so it can stop routing traffic to unhealthy servers.

Can a load balancer itself be a single point of failure?

Yes — which is why production systems often run multiple load balancers with their own failover mechanism.