Scalability
A system's ability to keep working well as usage grows — more users, more data, more requests.
What is it?
A system that works fine with a hundred users might completely fail with a million. Scalability is about designing a system so it can keep up as demand grows, ideally by adding more resources rather than needing a complete redesign.
There are two main strategies: vertical scaling (making a single server more powerful — more CPU, more memory) and horizontal scaling (adding more servers and spreading the work across them). Most large-scale systems eventually lean on horizontal scaling, since there's always a ceiling to how powerful one machine can get.
Explain like I'm 10
Vertical scaling is like hiring one super-employee who can work faster and faster. Horizontal scaling is like hiring more employees and splitting the work among them. At some point, no single employee — no matter how fast — can outpace hiring a team.
Examples
The idea, not literal code
// Vertical scaling: same one server, upgraded hardware
// 4 CPU cores, 8GB RAM → 32 CPU cores, 256GB RAM
// Horizontal scaling: same modest servers, more of them
// 1 server → 10 servers, behind a load balancerHow it works
Scalable systems are usually designed so that individual pieces (servers, in particular) don't hold irreplaceable state that only they know about — so any of them can handle any request, and more can be added freely. This often relies on other concepts working together: load balancers to distribute traffic, caches to reduce repeated work, and databases designed to handle growing data volumes.
Why does it exist?
User growth is often the whole point of building a successful product — but a system that can't scale becomes slow, unreliable, or simply falls over exactly when it matters most: when it's finally popular. Designing for scalability from the start avoids painful, risky rewrites later.
When to use it
Think about scalability deliberately once real growth is a realistic near-term possibility — when you're designing a system you expect to succeed and need to handle meaningfully more users, data, or traffic than it does today.
When not to use it
Don't over-invest in horizontal scaling, statelessness, and distributed architecture for a prototype or an app with a small, known, stable user base — that complexity has a real cost, and premature scaling work is a common way projects get bogged down before they even ship.
Common mistakes
Only ever scaling vertically, until hitting the hard ceiling of the most powerful single machine available.
Storing important state on an individual server (like data in memory) that horizontal scaling then breaks, since other servers don't have access to it.
Over-engineering for a scale the product doesn't remotely need yet, adding needless complexity too early.
Practice exercises
- Easy:
Explain, in your own words, the difference between vertical and horizontal scaling.
- Medium:
Describe a design decision that would make horizontal scaling harder, and how you'd avoid it.
- Hard:
Sketch (in words) how load balancing, caching, and a scalable database work together to let a system handle 100x more users.
Interview questions
What's the difference between vertical and horizontal scaling?
Vertical scaling adds more power to a single machine; horizontal scaling adds more machines and spreads the work across them.
Why do most large systems eventually favor horizontal scaling?
Because there's a physical and cost ceiling to how powerful a single machine can become, while adding more machines can, in principle, continue indefinitely.
Why is storing state only on one server a scalability problem?
Because other servers can't see that state, so requests must always be routed back to that specific server — breaking the flexibility that horizontal scaling relies on.