Caching

Keeping a copy of frequently-needed data somewhere much faster to access, so you don't redo expensive work every time.

What is it?

Some operations are expensive — a complex database query, a slow calculation, a request to another service far away. If the same result is needed again and again, redoing that expensive work every single time is wasteful. Caching means storing a copy of the result somewhere fast (often in memory) so future requests can just reuse it instead of recomputing it.

Explain like I'm 10

It's like keeping a jar of pre-made coffee in the fridge instead of brewing a fresh pot every single time someone wants a cup. It's faster to grab an existing cup — you just have to remember to refill the jar occasionally.

Examples

A simple cache in front of a slow lookup

const cache = new Map();

async function getUser(id) {
  if (cache.has(id)) {
    return cache.get(id); // fast — no database call
  }

  const user = await database.query("SELECT * FROM users WHERE id = ?", [id]);
  cache.set(id, user);
  return user;
}

How it works

Before doing expensive work, the system checks the cache first. If the data is there (a "cache hit"), it's returned immediately. If not (a "cache miss"), the expensive work runs, and the result is stored in the cache for next time. Cached data is usually also given an expiration time, so it doesn't become permanently stale.

Request comes in
       ↓
 Is it in the cache?
   ↓yes            ↓no
return cached    do expensive work
  value             ↓
                store result in cache
                     ↓
                return result

Why does it exist?

Caching dramatically reduces load on slow or expensive resources (like databases) and makes responses feel instant to users, at the cost of occasionally serving slightly outdated data — a trade-off that's usually well worth it for data that doesn't change every second.

When to use it

Reach for caching when the same expensive result is requested repeatedly and doesn't need to be perfectly fresh every single time — a product page, a popular search result, a computed report.

When not to use it

Don't cache data that must always be perfectly up to date and changes constantly — a live account balance mid-transaction, for instance — or if you do, keep the cache lifetime extremely short and invalidate it deliberately whenever the underlying data changes.

Common mistakes

  • Caching data that changes frequently without a short enough expiration, leading to users seeing stale information.

  • Forgetting to invalidate (clear) a cached value when the underlying data changes.

  • Caching sensitive or user-specific data in a shared cache without properly separating it per user.

Practice exercises

  1. Easy:

    Explain, in your own words, the difference between a 'cache hit' and a 'cache miss'.

  2. Medium:

    Add an expiration time to the cache example above, so cached values are only reused for 60 seconds.

  3. Hard:

    Describe a real scenario where caching stale data could cause a real user-facing problem, and how you'd mitigate it.

Interview questions

What is caching, in simple terms?

Storing a copy of a result somewhere fast to access, so repeated requests for the same thing don't redo expensive work.

What is cache invalidation, and why is it considered hard?

It's the process of removing or updating cached data once it's no longer accurate — hard because you must reliably track every place a cached value could become stale.

What's a trade-off caching introduces?

It can serve slightly outdated ('stale') data for a period of time, in exchange for much faster responses and less load on the underlying system.

What is the cache-aside (lazy-loading) pattern?

The application checks the cache first; on a miss, it reads from the underlying data source itself, then writes that result into the cache before returning it. The cache only ever holds what's actually been requested, and the application — not the cache — owns the read-through logic.

What is write-through caching, and how does it differ from cache-aside in when data enters the cache?

In write-through, every write goes to the cache and the underlying store together, as one operation, so the cache is always populated with the latest value the moment it's written. In cache-aside, data only enters the cache reactively, on a later read after a miss — a fresh write doesn't populate the cache by itself.

What is write-back (write-behind) caching, and what risk does it introduce that write-through avoids?

Write-back writes go to the cache immediately and are acknowledged as done, with the write to the underlying store happening asynchronously afterward. This is faster, but if the cache crashes before that async write completes, the data is lost — a risk write-through avoids by only confirming the write once both are durably updated.

What is "write-around" caching, and when would you use it?

Writes go directly to the underlying store, bypassing the cache entirely; the cache only gets populated later, on a read (like cache-aside). It suits data that's written once and rarely re-read soon after, avoiding filling the cache with values that may never be requested.

What is cache stampede (thundering herd), and give a concrete scenario where it happens.

When a popular cached key expires or is evicted, many concurrent requests can all miss at once and all hit the underlying data source simultaneously to recompute the same value — for example, a highly-requested product page's cache entry expiring right as a traffic spike hits, sending thousands of simultaneous queries to the database for the exact same row.

Name two mitigations for cache stampede.

A lock (or "single-flight") so only one request recomputes a missing value while others wait for that result instead of duplicating the work; or adding random jitter to TTLs so many entries don't all expire at the exact same instant.

What's the tradeoff in choosing a very short TTL versus a very long one?

A short TTL keeps data fresher but causes more frequent cache misses, pushing more load back onto the underlying system and reducing the benefit of caching at all. A long TTL maximizes cache hits and load reduction but risks serving stale data for longer after the underlying data changes.

What is "stale-while-revalidate," and what problem does it solve?

When a cached value expires, the request is still served the (now-stale) cached value immediately, while a background request refreshes it for next time — avoiding making the user's request wait on the slow recompute, at the cost of that one response being slightly outdated.

What is negative caching, and why cache the absence of something?

Caching the fact that a lookup found nothing (e.g. "no user with this id") so repeated requests for that same missing value don't repeatedly hit the expensive underlying system just to be told 'not found' again — useful when misses are frequent, like repeated lookups for invalid or deleted ids.

What's the difference between LRU and LFU eviction policies, and when would LFU beat LRU?

LRU (Least Recently Used) evicts whatever hasn't been accessed in the longest time; LFU (Least Frequently Used) evicts whatever has been accessed the fewest times overall. LFU can beat LRU when a genuinely popular item is only briefly not accessed (LRU might wrongly evict it for a rarely-used item that happened to be touched more recently).

What happens if a cache runs out of memory and no eviction policy is deliberately chosen?

Most cache systems fall back to some default eviction behavior (commonly LRU) rather than simply refusing new writes or crashing — but relying on an undchosen default means you haven't actually reasoned about which data you'd rather keep, which can silently evict something important under memory pressure.

What's the difference between a local (in-process) cache and a distributed cache like Redis, and what does each trade off?

A local cache lives in one application server's own memory — extremely fast with zero network hop, but only that one server sees it, and it disappears if the process restarts. A distributed cache is a separate shared service every server can read from — consistent across servers, but adds a network round trip and its own operational dependency.

Scenario: 10 app servers each keep their own local in-memory cache. What consistency problem does this create, and how would a distributed cache change it?

Each server's cache can hold a different, independently-stale version of the same key — a user could get a different answer depending on which server handles their request. A shared distributed cache gives every server the same view of cached data, at the cost of a network call each server must now make instead of a local memory read.

What is cache key design, and why does a poorly designed key silently break caching correctness?

The cache key must uniquely capture everything that affects the cached result — if two logically different requests (e.g. different filters, different users, different locales) end up mapping to the same key, one request can silently receive another's cached result, which is a correctness bug, not just a performance one.

Scenario: an API endpoint takes query params for pagination and filters. How would you design the cache key, and what happens if you get it wrong?

The key should encode every param that changes the response — page number, page size, and each filter value, typically normalized into a consistent order. Get it wrong (e.g. keying only on the endpoint path) and a request for page 2 could be served page 1's cached response, or one user's filtered results could leak to another user's differently-filtered request.

Why does adding a cache sometimes make debugging production issues harder?

A bug can now be intermittent and depend on cache state — the same request can behave differently on a hit versus a miss, or show a problem only after data changed but before the cache caught up, making it harder to reproduce and reason about than a system with one consistent code path.

What is cache warming, and when is it necessary?

Proactively populating a cache with expected data before real traffic arrives, rather than waiting for organic cache misses to fill it. Useful right after a deploy, cache restart, or expected traffic spike, so the system isn't hit with a wave of cold-cache misses all at once.

Scenario: a freshly deployed, empty distributed cache goes live right as a traffic surge hits. What could go wrong, and how would you prevent it?

Every request is effectively a cache miss at once, so the full weight of that surge lands directly on the underlying database or service exactly when it's least prepared for it — potentially overwhelming it. Cache warming beforehand, or a gradual traffic ramp-up, avoids sending a cold cache straight into peak load.

What's the difference between invalidating a cache entry and simply letting it expire via TTL?

Invalidation is an active, immediate removal or update triggered by a known change to the underlying data — the cache is corrected right away. TTL expiration is passive and time-based — the entry might still be served as stale for up to the full TTL window even after the underlying data has already changed, unless invalidation is also used.

Scenario: an e-commerce product's price is updated in the database, but the cache isn't explicitly invalidated. What could go wrong, and for how long?

Customers could keep seeing (and even purchasing at) the old cached price until the TTL naturally expires — potentially selling at a stale price for the entire TTL window, which is exactly why price-sensitive writes usually pair a short TTL with explicit invalidation on update, rather than relying on TTL alone.

What is "cache coherence," and why is it harder in a multi-server or multi-region setup?

Cache coherence is the property that every cache holding a copy of the same data agrees with the others. It gets harder across many servers or regions because invalidating one cache doesn't automatically invalidate the others — you need a way (like a pub/sub invalidation broadcast) to propagate the change everywhere a copy might exist, and that propagation itself takes time.

How does application-level caching differ from HTTP/CDN caching in where each sits in the request path?

CDN/HTTP caching sits in front of the application entirely, often geographically close to the user, and can serve a response without the request ever reaching your servers. Application-level caching sits inside your own infrastructure, avoiding expensive internal work (like a database query) but still requires the request to reach your server first.

Why can caching mask an underlying performance problem rather than truly solving it?

If a query is slow because of a missing index or a poor data model, caching hides the symptom for cache hits but does nothing to fix the actual slowness — every cache miss (and every uncached path) still pays the full cost, and the problem resurfaces immediately if the cache is cleared, overwhelmed, or bypassed.

What is a lock-based ("single-flight") approach to preventing stampede, and how does it work mechanically?

When a key misses, the first request to notice acquires a lock (or marker) saying "I'm already recomputing this," does the expensive work, then releases the lock and populates the cache. Other concurrent requests for the same key see the lock and either wait for that result or briefly serve a stale value, instead of every one of them independently redoing the same expensive work.

Scenario: a cache uses a 24-hour TTL and is also manually invalidated on writes. What can still go wrong if the invalidation has a race condition with a concurrent read?

If a read fetches from the (soon-to-be-stale) cache at nearly the same moment a write is invalidating and repopulating it, the read can win the race and re-cache the old value right after invalidation cleared it — leaving stale data cached again for up to the next full TTL window, since nothing else will trigger another invalidation until the next write.

Why is caching a poor fit for data with strict consistency requirements, like a bank balance mid-transaction — and what's an alternative if you must speed up reads there?

Any cache introduces a window where the cached value can lag the true value, which is unacceptable when a stale read could let someone act on money that isn't really there. A better alternative is optimizing the read path itself (better indexing, a dedicated fast replica kept synchronously consistent) rather than a cache that trades correctness for speed.

What's the relationship between cache hit ratio and total system load, and what does a "good" hit ratio depend on?

Each cache hit is one request the underlying system never has to handle, so a higher hit ratio directly reduces load on it — but what counts as "good" depends entirely on the workload's actual repeat-access pattern; a system where nearly every request is for unique data can't have a high hit ratio no matter how the cache is tuned.

Follow-up: your cache's hit ratio unexpectedly drops from 95% to 40%. What would you check?

Whether the cache was recently cleared or restarted (cold cache), whether TTLs were shortened or a deploy changed cache keys (so old and new requests no longer match existing entries), whether traffic patterns shifted toward less-repeated/unique requests, or whether the cache is evicting entries early due to memory pressure.

What is consistent hashing, and why does a distributed cache often use it when adding or removing nodes?

Consistent hashing maps both keys and cache nodes onto the same conceptual ring, so each key is owned by the nearest node going around it. Adding or removing a node only reshuffles the keys near that one node, instead of a naive hash(key) % nodeCount scheme, where changing the node count remaps almost every key at once and causes a mass cache miss.

Common misconception: "the cache is just an optimization, it can't be a source of bugs." Why is this wrong?

A cache introduces its own state (what's cached, for how long, under what key) that can diverge from the source of truth — stale reads, a bad cache key causing one user's data to leak into another's response, or a stampede overwhelming the backend are all real, cache-caused bugs, not just missed performance.

Scenario: two concurrent requests both miss the cache for the same key at the same moment, with no stampede protection. Walk through what happens under cache-aside.

Both requests see a miss, both independently go to the underlying data source and do the full expensive work, and both then write essentially the same result back into the cache — the second write is redundant, but the real cost is that the underlying system took two (or, under heavier concurrency, many more) hits for what should have been one.