Rate limiting is a security control, not a counter

Tue Sep 08 2026

Rate limiting is a security control, not a counter

So before naming five algorithms, ask the first question: what exactly are you trying to smooth, and whose requests are you counting?

So start with a public password-reset endpoint. It may trigger database work, email delivery, and account-enumeration risk. Authenticated callers may have stable account identities; anonymous callers may share an enterprise NAT address; a botnet may distribute requests across thousands of IPs.

So choose the limiter whose failure mode you can afford. This guide works from the resource and threat model to burst behavior, identity, distributed state, overload handling, and HTTP signaling.

In this article

Rate limiting starts with an identity and a failure budget

RFC 6585 defines 429 this way: “The 429 status code indicates that the user has sent too many requests in a given amount of time (‘rate limiting’).” The protocol gives you the status code and leaves the dangerous decisions to your design.

So choose the principal, the counting scope, and the failure behavior. HTTP 429 chooses none of them for you. The RFC explicitly says it does not define how an origin identifies a user or counts requests.

Write down three answers before selecting an algorithm:

  1. What resource or behavior are you protecting?
  2. Which bursts are legitimate?
  3. What identity is expensive for the attacker to multiply?

So for the password-reset endpoint, the expensive action may be sending reset messages or testing whether accounts exist. Per-account limits can protect authenticated identities or requested account identifiers. IP limits provide a separate signal for anonymous traffic. Compromised accounts remain a separate problem: an attacker who controls valid accounts can pass an account-key check.

But per-IP limiting is only a useful anonymous fallback. An enterprise can place many legitimate users behind one address, while a botnet can multiply source addresses cheaply.

And every request updates state: a counter, token balance, timestamp set, or virtual deadline. In a distributed deployment, accuracy, availability, storage, clock behavior, and consistency become part of the security design.

The five algorithms differ mainly in what they allow at the edges

So focus on the behavior each algorithm produces. Read the burst column first; that is where capacity and abuse risk meet.

AlgorithmStateBurst behaviorMain strengthMain failure mode
Fixed windowO(1)Up to 2× across a boundarySimplest implementationBoundary burst
Sliding-window logO(N) timestampsExact rolling limitExactnessMemory growth
Sliding-window counterO(1)Approximate rolling limitCompact compromiseEstimation drift
Token bucketO(1)Explicit burst capacityAverage rate plus burstsBurst may overload downstream
Leaky bucketO(1) virtual state; O(queue) with queued requestsSmooth outputControlled release rateQueue and rejection semantics

Security comes from the key and the failure path as much as from the algorithm. The table gives fixed windows their strongest argument—simplicity—and exposes their cost at the boundary.

Fixed windows trade precision for cheap state

Divide time into fixed intervals, such as one minute, and keep a counter for each key. A fixed-window counter needs O(1) state per key; Redis can implement it with an atomic counter-and-expiry operation.

For a limit of 100 requests per 60-second window, a 60-second rolling interval that straddles the fixed-window boundary can contain up to 200 requests: 100 just before the reset and 100 just after. That is defined behavior. It isn’t a mysterious implementation bug. Use fixed windows when simplicity matters and downstream capacity can absorb the edge burst.

Sliding windows keep more history or estimate it

A sliding-window log stores every request timestamp for a key. On each request, it removes timestamps older than the interval, counts what remains, and accepts or rejects the new request. The semantics are exact. The price is O(N) state.

High-frequency keys multiplied across many clients can make timestamp history expensive, even in a Redis sorted set. Use this when exact rolling behavior matters and the traffic volume and key cardinality make the history affordable.

A sliding-window counter keeps only the previous and current window counts. It estimates the rolling total by weighting the previous count:

estimate = previous × (1 − elapsed / window) + current

Here, elapsed is measured from the start of the current window, and elapsed and window must use the same units, such as milliseconds. Halfway through a one-minute window, half of the previous count contributes to the estimate.

The method uses O(1) state, but traffic clustered at the end of the previous window can make the estimate drift. Cloudflare describes the appeal of two numbers and simple arithmetic. A sliding-window counter is a practical compromise when fixed-window edges are too loose and exact timestamp history costs too much.

Token buckets admit bursts by design

A token bucket refills at rate r and holds at most b tokens. Each request consumes one token. Unused tokens create permission for an immediate burst.

The refill rate controls the long-term average; capacity controls the instantaneous burst. A bucket configured for 10 tokens per second and capacity 20 can admit 20 requests immediately when the bucket is full, assuming one token per request. After that, tokens return at 10 per second.

That model suits ordinary API reads and traffic spikes. Stripe uses token buckets for API rate limiting because the model keeps the average rate steady while allowing bursts.

Token bucket provides the same security as a sliding window only when its deliberate burst allowance fits the threat model. For password-reset or login-adjacent work, set capacity from downstream tolerance and abuse risk. A general API policy may allow too much.

Leaky buckets smooth release instead of admission

A token bucket controls admission. A leaky bucket controls the rate at which accepted work is released or processed. That distinction explains their different burst behavior.

A queue-based implementation stores queued work, while GCRA represents the schedule compactly; choose a tested implementation rather than assuming equivalent behavior.

nginx documents a request-rate limiter with rate, burst, and optional delay behavior. With rate=1r/s and burst=5, nginx allows five requests above the steady rate. Those requests join the permitted burst. With nodelay, nginx releases the burst immediately while still enforcing the configured rate over time. nginx rejects excess requests with 503 by default, rather than 429.

Choose from the downstream failure you can tolerate, not from the diagram you recognize.

Normal traffic exposes a trade-off the diagrams hide

A reproducible experiment from Yuhi-sa used these settings:

  • Window limit: 10 requests per second
  • Token bucket: 10 tokens per second, capacity 20
  • Leaky bucket: 10 requests per second, capacity 20
  • Traffic: Poisson arrivals at 6 requests per second
  • Duration: 30 seconds
  • Requests: 184
  • Random seed: 42

The results were:

  • Fixed window: 178 allowed, 6 rejected, or 96.74%
  • Sliding-window counter: 178 allowed, 6 rejected, or 96.74%
  • Token bucket: 184 allowed, 0 rejected, or 100%
  • Leaky bucket: 184 allowed, 0 rejected, or 100%

At 60% of the configured average rate, Poisson arrivals still cluster. Bucket capacity absorbs that jitter. Window methods reject some clustered requests because their counters see local peaks.

This demonstrates one workload, configuration, duration, and seed. Production capacity and user impact cannot be inferred from it; it says nothing about your storage latency, endpoint cost, or client retry pattern. Measure those separately.

The hardest bugs are usually about identity, overload, and semantics

“Why did my 100-per-minute limit allow 200?”

The fixed-window boundary did exactly what it was built to do. Accept that behavior when the limiter is a coarse fairness control and downstream capacity can absorb the edge. Use a sliding-window counter for constant-size state with less boundary distortion, or a log when exact history justifies its memory cost.

“Why doesn’t per-IP protection stop the attack?”

Per-IP limiting is a useful anonymous fallback for network-based traffic.

BackendBytes describes a 100-requests-per-minute fixed-window limiter keyed by IP that an attacker bypassed by spreading traffic across 1,000 IP addresses.

The opposite failure is just as real. A large enterprise may send many legitimate account holders through one NAT address, throttling unrelated users together. For authenticated traffic, key the primary limit to a stable account or user identity. For anonymous traffic, IP can be one signal.

For the password-reset endpoint, use independent account or requested-identifier limits alongside network-signal limits. That is usually safer than replacing both with one composite key: changing either component should not remove the other control. Test the policy against NAT collisions and distributed attacks before treating it as protective.

Account-key limits also have a boundary: a compromised account can make legitimate-looking requests. The limiter reduces one abuse path; it does not establish that the caller is benign.

“Why did the limiter increase the work during the attack?”

An application-layer server may still need to receive the request. It then runs the limiter, builds a response, and sends it. Under attack, generating a response for every rejected request adds work to the system you are trying to protect.

RFC 6585 permits dropping connections when responding to each request would consume too many resources. That does not make dropping connections universally preferable. Edge filtering, connection shedding, or upstream controls may need to act before application-level 429 generation.

See your DDoS protection and web application firewall controls as part of the same failure path, if those controls are present in your deployment.

“Why did my client retry the wrong thing?”

Because Retry-After is optional, clients must define a fallback. Save the actual retry procedure for the HTTP contract below.

“Why did the CDN serve a rate-limit response?”

RFC 6585 says 429 responses MUST NOT be cached. Verify that CDN and proxy layers preserve that rule, especially when responses contain account-specific or retry-specific details.

Distributed correctness costs more than adding Redis

The algorithms’ trade-offs are explainable, but safe limits cannot be chosen from theory alone. Measure endpoint cost, legitimate burst size, and limiter-store behavior under your traffic.

An in-memory counter works when one process makes every enforcement decision, or when per-instance limits are intentional. With multiple API instances, each process sees only part of the traffic unless state is shared or the policy is deliberately partitioned.

Use this sequence:

  1. Choose the key. Document the identity, fallback behavior, and trust boundary.
  2. Choose the state model. Counters, two-window estimates, timestamps, token balances, and virtual deadlines have different storage costs.
  3. Use shared state. Redis or an equivalent store keeps enforcement points from diverging.
  4. Make the decision atomic. The check and update must happen together. Redis Lua scripts are a common practical approach; atomic server-side operations or equivalent primitives can serve the same purpose.
  5. Define failure behavior. Decide what happens when the store is slow or unavailable: fail open, fail closed, or use a coarser emergency control. Also define expiry, clock handling, and regional behavior.
  6. Measure the limiter. Track decision latency, store errors, rejected requests, and key cardinality.

For leaky-bucket semantics, GCRA can represent a virtual schedule compactly. Use a tested implementation; the algorithm choice does not establish its performance or correctness for your store.

Clients need a rate-limit contract they can obey

A status code is part of the attack surface when the attacker controls how many rejected requests you process.

A server returning 429 should include a useful, non-sensitive explanation. RFC 6585 says the response should contain details and may include Retry-After.

Retry-After has two formats

The header can contain delta-seconds:

Retry-After: 3600

That means 3,600 seconds. It can also contain an HTTP date:

Retry-After: Wed, 21 Oct 2015 07:28:00 GMT

Clients must parse both formats, as documented in MDN’s Retry-After reference.

A client policy should:

  1. Honor a valid Retry-After.
  2. Otherwise use exponential backoff.
  3. Add full jitter.
  4. Cap the maximum delay.
  5. Stop retrying when the operation is no longer useful.

Many APIs expose de facto X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset fields. Draft IETF RateLimit-* fields are intended to standardize equivalent information, but deployment remains uneven. Document the fields your service actually supports.

nginx’s documented limit_req behavior emits 503 for rejected excess requests by default. Clients seeing 503 may retry as if the service is unavailable, while clients seeing 429 use rate-limit behavior. Test the response at the edge, not only in application code.

Pick the failure mode you can afford

  • Fixed window: choose the simplest possible control and accept its edge behavior.
  • Sliding-window log: choose exact rolling semantics when timestamp state is affordable.
  • Sliding-window counter: choose an O(1) approximation when fixed-window edges need correction.
  • Token bucket: choose a long-term average with legitimate bursts.
  • Leaky bucket: choose smooth processing and verify whether the implementation queues or rejects.

Before shipping, answer five questions:

  1. Is the key an identity or merely a network location?
  2. What burst does the algorithm permit?
  3. Where is shared state stored, and is the update atomic?
  4. What happens when the limiter or store is overloaded?
  5. Do server and client agree on status, Retry-After, headers, caching, and retries?

Log the key type, algorithm, decision, remaining allowance, store latency, and response status for a sampled request set. If you can’t explain a rejection, the limiter isn’t ready.