What is System Design?

Planning how the different parts of a real software system fit together and work at scale.

What is it?

Writing a single function or a small app is one skill. Deciding how dozens of services, databases, and servers should work together — so that millions of people can use a product reliably, quickly, and without it falling over — is a different skill entirely. That's system design.

It's less about writing code and more about making decisions: Where should data live? How do different parts of the system talk to each other? What happens when one part fails, or when ten times more people show up at once?

Explain like I'm 10

If writing code is like building a single room, system design is like being the architect of an entire city — deciding where the roads, water pipes, and power lines go so that everything works together, even as the population grows.

Examples

Questions system design answers

// Not code you'd run — these are the kinds of questions system design asks:

// - Should this data live in one database or be split across several?
// - What happens if this server crashes right now?
// - How do we serve 10 million users instead of 10 thousand?
// - Should two services talk directly, or through a queue?

How it works

System design usually starts with understanding requirements (how many users, how much data, how fast does it need to respond), then breaks the problem into components (servers, databases, caches, queues), and decides how those components connect and what happens when something goes wrong. It's an iterative process of trade-offs, not a single right answer.

Why does it exist?

A system that works perfectly for 10 users can completely fall apart at 10 million — slow responses, crashes, lost data. System design exists to think through those problems before they happen, and to make deliberate, informed trade-offs instead of accidental ones.

When to use it

Reach for system-design thinking whenever you're planning something bigger than a single script — choosing how services, data, and traffic should be organized before (or while) you build, especially once more than one person or more than a handful of users will depend on it.

When not to use it

For a small script, a personal tool, or a prototype nobody else depends on yet, heavy system design is overkill — you'd spend more time planning for scale you don't have than actually building. Start simple, and bring in system design as real constraints (more users, more data, more reliability needs) show up.

Common mistakes

  • Jumping straight to specific technologies before understanding the actual requirements and constraints.

  • Designing only for the current scale and ignoring how the system would need to change if usage grew 100x.

  • Assuming there's one 'correct' design instead of a best trade-off given the specific goals.

Practice exercises

  1. Easy:

    Pick an app you use daily (e.g. a messaging app) and list 3 questions you'd need to answer to design its backend.

  2. Medium:

    Sketch (in words) the major components you'd expect behind a simple photo-sharing app: what stores the photos, what stores the metadata, what serves requests.

  3. Hard:

    Describe what might break in your sketch above if the app suddenly went from 1,000 to 10,000,000 users.

Interview questions

What is system design, in your own words?

The process of deciding how the components of a software system (servers, databases, caches, queues, etc.) fit together to meet requirements like scale, speed, and reliability.

Why can't you just "add more servers" to fix every scaling problem?

Because bottlenecks often move to somewhere else — like a single database — that doesn't automatically scale just by adding more application servers.

What's the difference between a functional and a non-functional requirement in system design?

Functional requirements describe what the system does (e.g. 'users can post a photo'); non-functional requirements describe how well it does it (e.g. 'responses under 200ms', 'handles 1M users').

What's the difference between vertical scaling and horizontal scaling?

Vertical scaling means making a single machine more powerful (more CPU, RAM); horizontal scaling means adding more machines and spreading the load across them.

Why is horizontal scaling generally preferred for large-scale systems, despite being more complex?

A single machine has a hard ceiling on how powerful it can get and becomes a single point of failure, while horizontal scaling has no fixed ceiling and lets the system tolerate individual machines failing.

What are some of the main non-functional requirements a system designer needs to gather before designing a system?

Things like availability (how often it must be up), latency (how fast it must respond), scalability (how much growth it must handle), consistency (how up-to-date data must be across replicas), and durability (whether data can ever be lost).

What is a single point of failure, and why does system design try to eliminate them?

Any single component whose failure takes down the whole system — system design tries to remove them (through redundancy and replication) because a system is only as reliable as its least reliable required component.

Why do system design interviews emphasize trade-offs rather than a single correct answer?

Because real systems balance competing goals (cost, consistency, latency, complexity) differently depending on their specific requirements, so a good designer explains the reasoning behind a choice rather than reciting one 'correct' architecture.

What is back-of-the-envelope estimation, and why is it a common first step in system design?

Rough capacity math (expected users, requests per second, storage growth) done early to reveal which parts of a design will actually be under strain, so effort isn't wasted optimizing something that was never going to be a bottleneck.

What's the difference between throughput and latency?

Latency is how long a single request takes to complete; throughput is how many requests the system can handle in a given period — a system can have high throughput and still feel slow per-request, or vice versa.

Why might a design that works fine for 10,000 users fail at 10 million, even with no bugs?

Because resources that seemed effectively unlimited at small scale — a single database's connections, disk I/O, a single server's memory — become real bottlenecks once traffic and data volume grow by orders of magnitude.

What is the CAP theorem, and why does it matter when designing a distributed system?

It states that a distributed system can't simultaneously guarantee Consistency, Availability, and Partition tolerance during a network failure — since partitions are unavoidable in practice, it forces an explicit choice between staying consistent or staying available when the network splits.

What's the difference between designing for a read-heavy workload versus a write-heavy workload?

Read-heavy systems tend to lean on caching and read replicas to serve the same data to many readers cheaply; write-heavy systems need to focus on how writes are distributed and ordered (sharding, queues) since caching doesn't help data that changes constantly.

Why is it important to clarify assumptions and scope before diving into a system's design?

Without agreeing on scale, key features, and constraints up front, two people can design very different, equally 'correct' systems for the same vague prompt — clarifying scope focuses effort on the requirements that actually matter.

What does "availability" mean in system design, and how is it typically expressed?

The proportion of time a system is operational and able to serve requests, often expressed as a percentage of uptime (e.g. 99.9%, informally '3 nines') or in terms of allowed downtime per year.

Give an example of a trade-off between consistency and availability.

A banking system might reject a request during a network partition rather than risk showing a stale balance (favoring consistency), while a social media feed might show slightly outdated data rather than fail entirely (favoring availability).

Why might adding a cache introduce new problems instead of simply making a system faster?

A cache can serve stale data if it isn't invalidated correctly, adds an extra component that can fail or get out of sync with the source of truth, and shifts complexity toward cache-invalidation bugs, which are notoriously easy to get wrong.

What's the risk of over-engineering a system design for scale it doesn't actually need yet?

It burns time and money building and maintaining complexity (extra services, replication, queues) that isn't solving a real current problem, while likely guessing wrong about which future bottlenecks will actually matter.