Circuit Breaker & Retries
Protecting a system from a failing dependency, instead of letting that failure cascade everywhere.
What is it?
When one service calls another (common in microservices), that other service might be slow or down. Naively retrying the same failing call over and over can make things worse — it piles on more load exactly when the failing service can least handle it, and can even take down the caller too, as it piles up requests waiting on a response that never comes. A circuit breaker guards against this: after enough failures, it "trips" and stops sending requests to the failing dependency for a while, failing fast instead — giving the struggling service room to recover.
Explain like I'm 10
It's like an electrical circuit breaker in a house: rather than let a dangerous surge keep flowing and risk starting a fire, the breaker trips and cuts the circuit entirely. After some time (or once it's safe), it can be reset and power flows again.
Examples
A simple circuit breaker wrapping a risky call
let failureCount = 0;
let open = false;
let openedAt = null;
async function callWithBreaker(fn) {
if (open) {
if (Date.now() - openedAt < 30000) {
throw new Error("Circuit open — failing fast");
}
open = false; // try again after the cooldown
}
try {
const result = await fn();
failureCount = 0; // success resets the count
return result;
} catch (error) {
failureCount++;
if (failureCount >= 5) {
open = true;
openedAt = Date.now();
}
throw error;
}
}How it works
A circuit breaker tracks recent failures for a given dependency. While failures stay below a threshold, it behaves normally (closed — requests flow through). Once failures cross that threshold, it opens — further calls fail immediately, without even attempting the network call, for a cooldown period. After that cooldown, it allows a small number of test requests through (half-open) to check if the dependency has recovered, closing again if so.
Why does it exist?
Blindly retrying a failing dependency wastes time and resources on calls that are likely to fail anyway, and can pile up enough waiting requests to bring down the calling service too. A circuit breaker fails fast instead, protecting both the caller's own stability and giving the failing dependency breathing room to recover.
When to use it
Use a circuit breaker around calls to any external dependency (another microservice, a third-party API) whose failure shouldn't be allowed to cascade into your own service failing too — especially in systems with many service-to-service calls.
When not to use it
For calls where a failure genuinely should just fail immediately and be reported (no retrying makes sense, like a clearly invalid request), a circuit breaker adds complexity without benefit — it's specifically valuable for transient, recoverable failures.
Common mistakes
Retrying a failing call immediately and indefinitely, without any backoff or circuit breaker, piling on load exactly when the dependency is already struggling.
Setting the failure threshold or cooldown so aggressively that the breaker trips on ordinary, brief blips instead of genuine sustained failures.
Forgetting to have a sensible fallback behavior for when the circuit is open, instead of just surfacing a confusing error to the end user.
Practice exercises
- Easy:
Explain, in your own words, why blindly retrying a failing call can make an outage worse, not better.
- Medium:
Describe the three states of a circuit breaker (closed, open, half-open) and what triggers moving between them.
- Hard:
Explain what a good 'fallback' behavior might look like for a circuit breaker that's currently open, for a feature like showing product recommendations.
Interview questions
What problem does a circuit breaker solve?
It stops a caller from continuing to hammer a failing dependency with requests, failing fast instead to protect both the caller's own stability and give the dependency room to recover.
What are the three states of a circuit breaker?
Closed (requests flow normally), open (requests fail immediately without attempting the call), and half-open (a few test requests are allowed through to check if the dependency has recovered).
Why is blind, unlimited retrying dangerous?
It can pile up load on an already-struggling dependency and cause the caller itself to back up waiting on responses, potentially cascading the failure further.