Monitoring & Observability

Being able to tell what your system is actually doing, and why something went wrong, without guessing.

What is it?

A system running in production will eventually behave unexpectedly — slow down, throw errors, or fail outright — and when it does, someone needs to figure out why, often under time pressure. Monitoring means continuously tracking key signals about a system's health (like error rates, response times, and resource usage) and alerting when something looks wrong. Observability is the broader ability to actually understand why, using tools like logs (a record of discrete events), metrics (numeric measurements over time), and traces (the path a single request took across multiple services).

Explain like I'm 10

Monitoring is like a car's dashboard warning lights — it tells you something is wrong (check engine, low oil) without necessarily telling you exactly why. Observability is like being able to pop the hood and actually trace the problem back to its source — a specific wire, a specific part — using the detailed information available, rather than just knowing a light came on.

Examples

The three pillars, applied to one slow request

// A metric: how many requests are slow right now?
metrics.increment("api.requests.slow", { endpoint: "/checkout" });

// A log: what exactly happened for this one request?
logger.info("Checkout failed", { userId: 42, reason: "payment timeout" });

// A trace: which specific step, across multiple services, was slow?
// checkout-service (12ms) -> payment-service (4800ms) <- the slow one
//                         -> inventory-service (8ms)

How it works

Metrics are lightweight numeric counters and measurements collected continuously (requests per second, average latency, error rate) and are cheap to store and graph over time, making them great for spotting trends and triggering alerts. Logs capture detailed, timestamped records of specific events, useful for digging into exactly what happened. Traces follow one individual request as it moves through multiple services, showing exactly where time was spent — essential once a system is split across many services, where a single log or metric alone can't show the full path.

Why does it exist?

Without monitoring and observability, the only way to know something is wrong is when a user complains — and even then, there'd be no good way to figure out why. These tools let a team notice problems quickly (often before users do) and diagnose the actual root cause instead of guessing.

When to use it

Build in monitoring and observability from early on in any production system — metrics and alerts for overall health, logs for debugging specific incidents, and traces especially once a system involves multiple services calling each other.

When not to use it

For a small prototype or a personal project with no real users depending on it, full observability tooling (dedicated tracing infrastructure, alerting pipelines) is likely overkill — basic logging is often enough until the system actually matters to someone besides you.

Common mistakes

  • Logging so much low-value detail that genuinely important log entries get lost in the noise.

  • Having metrics and dashboards but no alerts, so problems are only noticed when someone happens to be looking at a graph.

  • Relying only on logs in a multi-service system, where a single request touches several services and no one log tells the full story — that's exactly what tracing is for.

Practice exercises

  1. Easy:

    Explain, in your own words, the difference between a metric and a log.

  2. Medium:

    Describe a scenario where a metric alone would tell you something is wrong, but you'd need a trace to find out why.

  3. Hard:

    Explain why observability becomes significantly more important once a system moves from a monolith to microservices.

Interview questions

What's the difference between monitoring and observability?

Monitoring tracks key signals and alerts when something looks wrong; observability is the broader ability to actually understand why, using logs, metrics, and traces together.

What are the three commonly cited 'pillars' of observability?

Logs (detailed records of specific events), metrics (numeric measurements over time), and traces (the path of a single request across multiple services).

Why does tracing become especially important in a microservices architecture?

Because a single request can pass through many separate services, and no single log or metric can show the complete path — a trace ties all the pieces together.