Microservices vs Monolith

Two different ways to structure an application's codebase and deployment — as one unit, or as many independent pieces.

What is it?

Early in a project, it's simplest to build everything — the user system, the payments, the notifications — as one single application, deployed and scaled as a single unit. This is a monolith. As a system and its team grow, some organizations split that single application into many smaller, independently deployable services, each responsible for one part of the system, communicating over the network — this is a microservices architecture.

Neither is universally correct: a monolith is simpler to build, test, and deploy; microservices offer independent scaling and deployment, at the cost of real operational complexity.

Explain like I'm 10

A monolith is like one large, all-purpose kitchen where every dish is prepared by the same staff in the same space. Microservices are like a food court, where each stall specializes in one thing and operates independently — more flexible and easier to scale one popular stall without touching the others, but now you need to coordinate delivery and communication between many separate operations instead of one.

Examples

The same feature, two different structures

// Monolith: one codebase, one deployment
function handleOrder(order) {
  chargePayment(order);   // same process
  updateInventory(order); // same process
  sendConfirmationEmail(order); // same process
}

// Microservices: separate services, communicating over the network
async function handleOrderMicroservices(order) {
  await paymentService.charge(order);       // a network call
  await inventoryService.update(order);     // another service
  await notificationService.notify(order);  // yet another
}

How it works

In a monolith, every part of the application runs as one process, sharing memory and typically one database, so calling between parts is just a regular function call. In microservices, each service is its own separate process (often with its own database), communicating with other services over the network — usually via HTTP APIs or message queues — which introduces network latency, partial failure (one service can be down while others work), and the need for careful API contracts between services.

Why does it exist?

As a codebase and a team both grow large, a monolith can become hard to work on — every change risks affecting unrelated parts, and the whole thing must be deployed together. Microservices let different teams own, deploy, and scale their own piece of the system independently, at the cost of significant added operational complexity.

When to use it

Start with a monolith for most new projects — it's simpler to build, easier to reason about, and faster to change early on. Consider splitting into microservices once a team has genuinely grown large enough that independent teams need to deploy independently, or specific parts of the system have wildly different scaling needs.

When not to use it

Don't reach for microservices just because large companies use them — for a small team or an early-stage product, the operational overhead (network calls, service discovery, distributed debugging) usually costs far more than it's worth. Many successful products run as a monolith for years.

Common mistakes

  • Adopting microservices prematurely, taking on distributed-systems complexity before the team or product actually needs it.

  • Splitting services along lines that don't match how the team is organized, so every feature still requires coordinating across many services anyway.

  • Forgetting that a network call between microservices can fail in ways a normal function call never could, and not handling those failures.

Practice exercises

  1. Easy:

    Explain, in your own words, the main tradeoff between a monolith and microservices.

  2. Medium:

    Describe a warning sign that a monolith might be becoming difficult for a growing team to work in.

  3. Hard:

    Explain why a single feature (like placing an order) becomes harder to reason about when it's split across three separate microservices instead of one monolith.

Interview questions

What is the main difference between a monolith and microservices?

A monolith is one deployable application containing all functionality; microservices split functionality into many independently deployable services that communicate over the network.

What's a real cost of microservices that a monolith doesn't have?

Network calls between services can fail independently, add latency, and require careful handling of partial failures — problems a single in-process monolith never encounters.

Why might a company choose microservices despite the added complexity?

To let independent teams deploy and scale their own services without coordinating with every other team, and to scale only the specific parts of the system that need it.

What is a "distributed monolith," and how does a team end up with one despite having "microservices"?

It's a system split into separate deployable services that still can't actually be deployed or changed independently — because they share a database, call each other synchronously in tightly-coupled chains, or share versioned libraries that force them to be updated in lockstep. It happens when a team splits the code without splitting the actual coupling, ending up with all of microservices' operational cost and none of its independence benefit.

What are some warning signs that you actually have a distributed monolith rather than true microservices?

You can't deploy one service without also deploying several others; a single feature routinely requires coordinated changes across many services at once; services share a database or shared mutable state; and a failure or slowdown in one service reliably cascades into failures in several others rather than degrading gracefully.

What does "each microservice owns its own data" mean, and why is a shared database across microservices considered an anti-pattern?

It means each service is the only one allowed to directly read or write its own underlying storage — other services must go through its API, never its database. A shared database re-couples services at the data layer even if their code is separate: one service's schema change can silently break another, and it becomes impossible to reason about which service is actually responsible for a given piece of data's integrity.

Scenario: an orders service and an inventory service share the same underlying database. What problems does this cause?

Either service can change the shared schema in a way that breaks the other without either team necessarily realizing it, neither can be deployed or scaled independently since both depend on one shared database's availability and schema version, and it becomes unclear whose responsibility it is to maintain data integrity — defeating a core reason to split them into separate services in the first place.

What is Conway's Law, and how does it argue for aligning service boundaries with team boundaries?

Conway's Law observes that a system's architecture tends to mirror the communication structure of the organization that built it. It argues you should draw service boundaries around how your teams are actually organized (so a service can be owned and changed by one team without needing constant cross-team coordination), rather than drawing boundaries first and hoping teams naturally reorganize around them.

What is a "bounded context" (from domain-driven design), and why is it a common way to decide where to draw a microservice boundary?

A bounded context is a boundary within which a specific business concept (like "order" or "inventory") has one consistent, well-defined meaning and model — outside that boundary, the same word might mean something subtly different to a different part of the business. It's a natural fit for service boundaries because each bounded context tends to have its own data, its own rules, and can reasonably be owned by one team.

What is the "monolith-first" approach, and why do many experienced teams recommend it even when they expect to eventually need microservices?

It means deliberately starting a new system as a single, well-structured monolith, and only splitting out services once real usage has revealed where the actual boundaries and scaling needs are. It's recommended because the right service boundaries are hard to guess correctly upfront, and a monolith is far cheaper to restructure internally than a set of prematurely-drawn microservice boundaries are to undo.

What is the strangler fig pattern, and how does it let you migrate a monolith to microservices incrementally?

New functionality (or a piece of existing functionality) is built as a new service that gradually intercepts and takes over a specific responsibility, while the monolith keeps running and handling everything not yet migrated — until, piece by piece, the monolith's remaining responsibilities shrink and it can eventually be retired. It avoids the risk of a single big-bang rewrite by migrating one bounded piece at a time, each independently verifiable.

What is the saga pattern, and what problem does it solve for a "transaction" that spans multiple microservices?

A saga breaks a multi-step business operation that spans several services into a sequence of local transactions, each with a corresponding compensating action that can undo it if a later step fails — since there's no single shared database to wrap the whole thing in one ACID transaction, a saga achieves an equivalent all-or-nothing outcome through explicit forward and backward steps instead.

Scenario: placing an order charges payment, reserves inventory, then schedules shipping, across three services with no shared database. Payment succeeds but inventory reservation fails — walk through how a saga handles this.

The saga recognizes the failure at the inventory step and triggers the compensating action for every step that already succeeded — in this case, refunding the payment that was already charged — rather than leaving the customer charged for an order that can never actually be fulfilled. The shipping step is simply never reached.

What's the difference between choreography-based and orchestration-based sagas?

In choreography, each service reacts to events from the others and decides its own next action independently, with no central coordinator — more decoupled, but harder to see the overall flow in one place. In orchestration, a central coordinator explicitly tells each service what to do and in what order, and handles triggering compensations — easier to reason about and observe, but that coordinator becomes a piece of shared logic every service in the saga depends on.

Why does distributed tracing become necessary in microservices in a way it isn't for a monolith?

In a monolith, a slow or failing request can be diagnosed with a single stack trace or profiler within one process. In microservices, a single user request can fan out across many independent services, each with its own logs, and there's no single process to inspect — distributed tracing stitches together the full cross-service path of one request (with per-hop timing) into one coherent view that would otherwise require manually correlating scattered logs from many different systems.

Scenario: a single user-facing request touches 6 services and is slow. Without distributed tracing, why is this hard to debug?

Each service only has visibility into its own portion of the work — none of them individually knows how long the other 5 took, or which one was actually the bottleneck. Without a shared trace id linking all 6 services' logs for that one request together, engineers are left manually cross-referencing timestamps across six separate logging systems to even figure out where the time went.

What is a circuit breaker, and why does calling another microservice over the network need one when calling a function within a monolith doesn't?

A circuit breaker detects that calls to a downstream service are consistently failing or timing out, and starts failing fast (without even attempting the call) for a cooldown period, rather than letting every caller keep waiting on a service that's already struggling. An in-process function call in a monolith either succeeds, throws, or the whole process crashes together — there's no independent "the other side is down but I'm still up" state a circuit breaker needs to protect against, because there's no network between them.

Why is versioning API contracts between microservices a much bigger deal than versioning function signatures within a monolith?

Within a monolith, changing a function's signature and updating every caller happens together, in the same deploy, checked by the same compiler/type system. Across microservices, the caller and the service can be deployed independently and at different times, so a breaking API change can be live in production being called by callers still expecting the old shape — contracts need explicit versioning and backward compatibility, since you can't guarantee every caller upgrades in lockstep.

What is contract testing, and what problem does it solve between independently-deployed services?

Contract testing verifies that a service's actual API matches what its consumers expect, without needing to spin up every real consumer and provider together to test it. It catches a breaking change to a service's contract before it's deployed, without requiring a full, slow, flaky end-to-end test across every dependent service just to know whether an integration still works.

Why is testing generally harder in a microservices architecture than in a monolith?

A monolith's tests can exercise a whole business flow in-process, fast and deterministically. In microservices, the same flow spans multiple independently-deployed services communicating over the network, so a true end-to-end test requires standing up several real (or realistically faked) services together, and network-related flakiness and version-mismatch bugs become possible in a way an in-process call never has to account for.

What does "independent deployability" actually buy a team, concretely — and what does it cost operationally?

It lets one team ship a change to their service without waiting for, or coordinating a simultaneous release with, every other team — faster, lower-risk releases scoped to just their own code. It costs real operational overhead: each service needs its own deployment pipeline, monitoring, and on-call ownership, and the team must design every API change to stay compatible with whatever version of every consumer happens to still be running.

Scenario: the payments team wants to deploy a breaking API change. What does true independent deployability require of how they roll this out?

They can't simply break the old contract in place — they need to either support both the old and new API shape simultaneously for a transition period (versioning), or coordinate the rollout with every known consumer to migrate first, precisely because other services may still be running against the old contract and deploy on their own separate schedule.

What is the "premature decomposition" trap, and why does it commonly happen to teams inspired by big tech blog posts?

It's adopting microservices because a well-known, much larger company uses them successfully, without having that company's actual scale, team size, or specific bottlenecks that justified the decision for them. The blog post shows the payoff, not the years of pain and the specific organizational scale that made the tradeoff worth it for that company at that size.

What's a realistic signal that a monolith has actually grown to the point microservices would help, rather than just "it feels big"?

Independent teams are genuinely blocked from deploying their own changes without coordinating releases with unrelated teams, or specific parts of the system have such different scaling/reliability needs that scaling the whole monolith to satisfy one part wastes significant resources on the rest — a felt sense of size or code sprawl alone isn't the same signal, and can often be fixed by better internal modularity instead.

Why can splitting a monolith along the wrong boundaries make things worse rather than better?

If a boundary is drawn through the middle of what's actually one cohesive piece of business logic, most real features end up needing coordinated changes across the resulting services anyway — so you now pay for network calls, versioning, and distributed debugging, without actually gaining the independent-deployment benefit the split was supposed to provide.

What operational capabilities does a team typically need in place before microservices become net-positive, that a monolith doesn't require?

Centralized logging and distributed tracing, service discovery, automated per-service deployment pipelines, monitoring/alerting per service, and generally some form of container orchestration — without this tooling in place, teams end up manually reproducing what a monolith gets essentially for free, one struggling service at a time.

Why does network partial failure require different error-handling thinking than a monolith's function calls?

A monolith's function calls either complete or the whole process crashes together — there's no state where one part of your own process is unreachable while the rest keeps running. Across a network, one dependency can be slow, timing out, or fully down while everything else stays healthy, so microservices code has to explicitly handle timeouts, retries, and partial degradation as ordinary, expected conditions rather than rare edge cases.

What is a service mesh, and what problems does it solve that would otherwise be re-implemented independently in every microservice?

A service mesh is an infrastructure layer (typically sidecar proxies alongside each service) that handles cross-cutting network concerns — retries, timeouts, circuit breaking, mutual TLS, load balancing, and observability — centrally and consistently, rather than every individual service reimplementing its own version of this logic in whatever language or framework it happens to use.

Scenario: after splitting into microservices, feature development actually slowed down because most features still touch three or four services at once. What does this suggest about how the boundaries were drawn?

It suggests the services were split along lines that don't match how the business logic actually varies — the boundaries likely followed a technical layering (e.g. one service per database table) rather than a cohesive business capability, so a single feature's logic ends up spread across several services that all still need to change together, paying microservices' coordination cost with none of its independence benefit.

Why is "our team is small" often a stronger argument against microservices than any specific technical concern?

Microservices' main benefit is letting multiple independent teams work and deploy without stepping on each other — a benefit that doesn't exist if there's really only one team, since that team still has to coordinate across every service they own regardless of how the code is split, while paying the full network, deployment, and observability overhead of the split anyway.

What's the relationship between eventual consistency and microservices — why do cross-service reads often end up eventually consistent even if each service's own database is strongly consistent?

Each service can be perfectly strongly consistent within its own database, but keeping one service's view of another service's data up to date (e.g. via events or async replication) happens over the network, on its own schedule — so from the perspective of a consumer relying on another service's data, there's inherently a window where that copy can be behind, just as with database replication, but now at the level of whole services instead of database replicas.

Common misconception: "microservices make a system more reliable." Why is this not automatically true?

Splitting a monolith into many services multiplies the number of independent things that can fail (each service, each network hop between them) — without deliberate resilience patterns (circuit breakers, retries with backoff, graceful degradation), a microservices system can actually be less reliable than a monolith, since a failure in any one of many dependencies can cascade. It becomes more reliable only when the added failure surface is actively engineered around, not as an automatic side effect of the split.

What's a realistic warning sign in a monolith that one specific piece (not the whole system) is a good first candidate to extract into its own service?

A piece of functionality with genuinely different scaling needs than the rest of the system (e.g. a heavy image-processing job dragging down an otherwise fast API), or one with a clearly separable, stable API boundary and its own data that a specific team could own end-to-end — a good extraction candidate is narrow and well-isolated, not "let's split everything at once."