Rate Limiting

Deliberately capping how many requests a client can make in a given time window.

What is it?

Without any limit, a single client — whether malicious, buggy, or just very active — could send an overwhelming number of requests and degrade the service for everyone else. Rate limiting is a deliberate rule that caps how many requests a given client (by user, API key, or IP address) can make within a certain time window, rejecting or delaying requests beyond that.

Explain like I'm 10

It's like a nightclub bouncer who only lets in a certain number of people per minute, no matter how many are waiting outside — not because the club dislikes visitors, but because letting everyone in at once would be a disaster for everyone already inside.

Examples

A simple fixed-window rate limiter

const requestCounts = new Map(); // key: userId, value: { count, windowStart }

function isAllowed(userId, limit = 100, windowMs = 60000) {
  const now = Date.now();
  const entry = requestCounts.get(userId);

  if (!entry || now - entry.windowStart > windowMs) {
    requestCounts.set(userId, { count: 1, windowStart: now });
    return true;
  }

  if (entry.count >= limit) return false; // over the limit
  entry.count++;
  return true;
}

How it works

Each incoming request is checked against a counter tied to that specific client, tracked over a time window (fixed windows, like "per minute", or smoother sliding windows). If the client is under their limit, the counter increments and the request proceeds; if they're at or over it, the request is rejected — commonly with an HTTP 429 "Too Many Requests" response — until the window resets.

Why does it exist?

Rate limiting protects a service from being overwhelmed — whether by a genuine traffic spike, a buggy client stuck in a retry loop, or a deliberate abuse attempt — and ensures one client's excessive usage can't degrade the experience for everyone else.

When to use it

Add rate limiting to any public-facing API, especially ones with expensive operations (search, sending emails, calling a paid third-party service) or ones that could be abused (login attempts, password resets).

When not to use it

Don't rate-limit so aggressively that legitimate, normal usage gets rejected — limits should be set based on real usage patterns, and it's often worth returning clear information (like how long until the limit resets) so well-behaved clients can adapt.

Common mistakes

  • Setting limits so low that normal, legitimate usage gets rejected, frustrating real users.

  • Rate limiting only by IP address, which breaks down for many users sharing one IP (like an office) and is easy to evade with many IPs.

  • Forgetting to tell the client why their request was rejected and when they can try again, making the limit feel arbitrary and confusing.

Practice exercises

  1. Easy:

    Explain, in your own words, why an API without any rate limit is vulnerable to a single misbehaving client.

  2. Medium:

    Describe the difference between rate-limiting by IP address versus by an authenticated user id, and a downside of each.

  3. Hard:

    Explain why a 'fixed window' rate limiter can allow twice the intended limit right at the boundary between two windows, and how a 'sliding window' avoids that.

Interview questions

What problem does rate limiting solve?

It prevents a single client from overwhelming a service with too many requests, whether from abuse, bugs, or unexpectedly high legitimate usage.

What HTTP status code typically indicates a rate limit was hit?

429 Too Many Requests.

What's a weakness of rate limiting purely by IP address?

Multiple legitimate users can share one IP address (like an office network), and it's relatively easy for an attacker to spread requests across many IPs to evade the limit.

What is the token bucket algorithm, mechanically?

A bucket holds tokens, refilled at a steady rate up to some maximum capacity; each request consumes one token to proceed, and is rejected (or queued) if the bucket is empty. It naturally allows short bursts up to the bucket's capacity while still enforcing a steady average rate over time via the refill rate.

What does the token bucket's capacity (burst size) control, separately from its refill rate?

The refill rate controls the long-run average allowed rate; the bucket's capacity controls how large a burst can be spent all at once before the client has to wait for tokens to trickle back in — two independent knobs, one for sustained rate and one for burst tolerance.

What is the leaky bucket algorithm, and how does it differ from token bucket in the traffic pattern it produces?

Requests enter a queue (the bucket) and are processed (leaked out) at a strictly constant rate, regardless of how bursty the incoming requests were. Unlike token bucket, which lets a burst through immediately up to its capacity, leaky bucket smooths bursts into a steady, constant output rate — good for protecting a downstream system that needs an even load, not just an average one.

Scenario: an API should allow occasional bursts but enforce a steady average rate. Token bucket or leaky bucket, and why?

Token bucket — its whole design is allowing accumulated capacity to be spent in a burst while still capping the long-run average via the refill rate. Leaky bucket would instead flatten that same burst into a steady trickle, which is the wrong fit if genuine bursts are supposed to be allowed through immediately.

What's the fixed window counter's boundary problem? Walk through it with a concrete example.

With a 100-requests-per-minute fixed window, a client could send 100 requests in the last second of one window (say, 11:00:59) and another 100 in the first second of the next window (11:01:00) — 200 requests in about 2 seconds, even though each individual window's limit was technically respected. The fixed window doesn't look at any sliding one-minute span, only at aligned clock-minute buckets.

What is the sliding window log algorithm, and why is it the most accurate but least memory-efficient approach?

It stores a timestamp for every individual request within the current window, and on each new request, counts how many stored timestamps fall within the last window-length of time, discarding older ones. It's precisely accurate because it looks at a true rolling window rather than aligned buckets, but it needs to store a timestamp per request rather than a single counter, which gets memory-expensive at high request volume.

What is the sliding window counter, and how does it approximate the log algorithm cheaply?

It keeps just two fixed-window counters (the current window and the previous one), and estimates the count within the true sliding window by weighting the previous window's count proportionally to how much of it still overlaps the sliding window — approximating the log's accuracy using only two numbers instead of a full list of timestamps.

Walk through the sliding window counter's formula with a concrete example.

If we're 25% into the current 1-minute window (15 seconds in), the estimate is: current window count + previous window count × (1 − 0.25), i.e. weighting 75% of the previous window's requests as if they still count toward the present sliding view — a close approximation of the true rolling count without storing every timestamp.

Why is distributed rate limiting harder than single-server rate limiting?

A single server can just keep a counter in its own memory, but if traffic for the same client is spread across many servers, each server's local counter only sees a fraction of that client's total requests — none of them individually knows the true total, so the limit can be silently multiplied by however many servers happen to handle that client's traffic.

Scenario: an API is rate-limited per-user, but requests are load-balanced across 10 stateless app servers, each keeping counts in local memory. What goes wrong?

A user's requests get spread roughly evenly across the 10 servers, and each server only counts what it personally saw — so a user could effectively make up to 10× the intended limit by having their requests spread across all 10 servers, none of which individually appears to exceed the per-server share of the limit.

What's the standard fix for distributed rate limiting?

Move the counter out of any individual server's local memory and into a single shared, fast store (commonly Redis) that every server checks and updates — so the count reflects a client's true total across every server, not just whichever one happened to handle a given request.

Why must the increment-and-check operation in a shared Redis-based rate limiter be atomic, and what race condition happens if it's not?

If two servers both read the current count, both see it's under the limit, and both then separately increment it, both requests get allowed even though the combined result now exceeds the limit — a classic check-then-act race condition. Making the read-check-increment one atomic operation closes that window entirely.

What Redis feature is commonly used to make this atomic?

Either Redis's INCR (which atomically increments and returns the new value in one operation, so the check can be safely done on the returned value) paired with EXPIRE for the window, or a small Lua script executed atomically inside Redis so the whole check-and-increment logic runs as a single indivisible step.

What impact does clock skew have on distributed rate limiting across multiple servers or regions?

If different servers disagree slightly on the current time, window boundaries can shift depending on which server's clock is used, letting a client's requests straddle what one server considers a fresh window while another still considers it the old one — effectively reopening a small amount of extra allowance. Using a single shared store's own clock (rather than each server's local clock) for window calculations avoids this.

What's the difference between rate limiting and backpressure/load shedding?

Rate limiting caps how much a specific client is allowed to send, based on a policy decided in advance, independent of the server's current load. Backpressure/load shedding is the server reactively slowing down or rejecting requests based on its own real-time capacity, regardless of which client they're from — a rate limit can still let requests through that overwhelm an already-struggling server if it doesn't happen to reflect actual current load.

Why might you rate limit by API key or user id rather than IP address for an authenticated API, and when would you still also want IP-based limiting?

An authenticated identity (API key/user id) is stable and specific to one actual client, avoiding the false sharing/evasion problems IP-based limiting has. You'd still also want IP-based limiting in front of authentication itself — for endpoints like login, where there's no user id yet to key off of, and you need to limit unauthenticated abuse before identity is even established.

What should a 429 response ideally include, and why does that matter for well-behaved clients?

It should tell the client how long to wait before retrying (commonly via a Retry-After header) and ideally how much of their quota remains and when it resets. Without this, a well-behaved client has no principled way to know when to retry, and may either hammer the API immediately (making things worse) or back off far more conservatively than necessary.

What is the `Retry-After` header, and how should a client be expected to use it?

It tells the client, in seconds or as a timestamp, when it's safe to retry the request. A well-behaved client should wait at least that long before retrying rather than retrying immediately or on its own arbitrary schedule, which helps the client recover gracefully without adding to the very overload that triggered the limit.

Scenario: a legitimate client (e.g. a batch job clearing a backlog) gets rate-limited during a real, valid traffic spike. What's a better strategy than a flat hard reject?

Consider a burst-tolerant algorithm (token bucket) that allows a temporary spike above the steady-state rate, or a separate, higher-throughput tier/queue for known bulk/batch clients — rather than uniformly capping every client at the same rate designed around typical interactive usage, which unfairly penalizes a legitimate, different traffic pattern.

What is a "leaky bucket as a queue" implementation, and how does it differ from leaky bucket purely as a rate limiter?

As a rate limiter, excess requests beyond the bucket's draining rate are simply rejected. Implemented as a queue instead, excess requests are held and processed later at the steady drain rate rather than dropped outright — trading immediate rejection for added latency, useful when a request shouldn't be lost, only delayed.

Why doesn't rate limiting alone fully protect against a DDoS attack distributed across many different IPs or accounts?

Rate limiting caps each individual client's allowed rate, but if an attack comes from thousands of different IPs or fake accounts, each one can stay comfortably under its own individual limit while the combined total traffic still overwhelms the service — rate limiting bounds per-client abuse, not aggregate volume from many coordinated sources, which needs additional defenses (traffic filtering, anomaly detection, a CDN/DDoS mitigation layer).

What layers of a system might apply rate limiting, and why might you want it at more than one?

It can be applied at the client SDK (self-throttling before even sending a request), an API gateway (a single, centralized enforcement point for external traffic), and individual services (protecting a specific expensive operation regardless of how a request arrived). Layering it catches different failure modes — a well-behaved client that self-throttles reduces unnecessary traffic entirely, while gateway and service-level limits protect against clients that don't.

Scenario: login attempts are rate-limited per account to prevent brute-force password guessing. What's a subtlety if you naively lock the account out after N failed attempts?

An attacker who knows this can deliberately trigger failed logins on a victim's account to lock the real user out of their own account — a denial-of-service against the legitimate user via the very protection meant to help them. Mitigations include increasing delay rather than a hard lockout, rate-limiting by source (IP/device) in addition to account, or requiring an additional verification step instead of fully denying login.

What's the tradeoff between rate limiting enforced on the server versus asking the client to self-throttle?

Server-side enforcement is authoritative and can't be bypassed by a misbehaving or buggy client, but every rejected request still costs the server some resources to receive and reject. Client-side self-throttling avoids sending excess requests at all, saving that cost, but can't be trusted alone since any client can simply ignore it — in practice, server-side enforcement is the actual guarantee, and client self-throttling is an optimization on top of it.

Why is choosing the rate-limiting window size (1 second vs 1 minute vs 1 hour) a real design decision, not an arbitrary one?

A shorter window reacts to abuse quickly but allows less flexibility for legitimate bursty usage within it; a longer window smooths out legitimate short bursts but lets a genuinely abusive client sustain a higher instantaneous rate for longer before the limit catches up to them — the right window size depends on what burst pattern is normal for real usage versus what pattern would actually indicate abuse.

What happens to a token bucket's burst allowance if a client is idle for a long period and then suddenly sends a large batch?

Tokens accumulate up to the bucket's maximum capacity while the client is idle, so the client can then spend that entire accumulated capacity as one large burst the instant it starts sending again — by design, but worth being deliberate about, since a very large bucket capacity paired with long idle periods can allow bursts far larger than the steady-state rate would suggest.

Common misconception: "rate limiting is only about stopping malicious abuse." Why is this incomplete?

Rate limiting just as often protects a service from its own legitimate clients — a buggy retry loop, a misconfigured cron job, or simply more real, valid usage than a downstream dependency (like a third-party API with its own limits) can handle. Framing it purely as an anti-abuse tool misses that it's also basic operational self-protection against ordinary bugs and success.

Scenario: an internal microservice calls another at a high, bursty rate that occasionally exceeds what the downstream service can handle, causing cascading failures. Is client-imposed rate limiting or something else the better fix?

Client-side self-throttling helps but relies on every caller cooperating and knowing the right limit in advance. A more robust fix is usually the downstream service itself enforcing its own rate limit or applying backpressure based on its real, current capacity — combined with the caller using a circuit breaker so it stops hammering a struggling downstream service and fails fast instead of piling up cascading retries.

Follow-up: how does rate limiting relate to a circuit breaker — are they solving the same problem?

Related, but different: rate limiting proactively caps how much traffic is sent based on a policy, regardless of whether the downstream is currently healthy. A circuit breaker reactively detects that a downstream is already failing and stops sending it traffic until it recovers — rate limiting prevents overload from happening, a circuit breaker responds once something is already going wrong.

Why can a sliding window log's memory usage become an operational concern at high request volume, and what does the sliding window counter trade for cheaper memory?

The log has to retain one timestamp per request within the entire current window, so at very high request rates (e.g. thousands of requests per second per key), that's a genuinely large amount of per-client memory. The sliding window counter trades exactness for just two integers per client (current and previous window counts), accepting a close approximation of the true rolling count instead of a fully precise one, in exchange for dramatically lower memory use.