Rate limiting looks like a solved problem until you run more than one instance of the thing doing it. Then the limit you configured stops being the limit anyone experiences, because each replica counts independently and the effective ceiling is your configured value multiplied by however many replicas exist right now. That number changes when you autoscale.
The fix is a shared counter, and the shared counter is where all the interesting problems live. Every request now involves a round trip to somewhere else, which is latency you added to your hot path. The counter is a new dependency, which is availability you gave away. And the moment that dependency is slow or unreachable, you have to decide something genuinely hard: do you let everything through and risk the overload you were protecting against, or do you reject everything and cause the outage yourself.
There is also a question nobody asks early enough, which is what the limit is for. Protecting a fragile backend from overload, enforcing a commercial quota on a paying customer, and stopping credential stuffing are three different problems with three different correct algorithms, three different accuracy requirements and three different failure behaviours. Teams that implement one mechanism for all three end up with a system that is too strict for the quota case and too loose for the abuse case.
So the algorithms come first here, then the distributed accuracy problem, then failure behaviour, and then the products that implement all of it.
Key takeaways
- Node-local rate limiting is not rate limiting once you have more than one replica. The effective limit is your configured value times the replica count, and it changes when you scale.
- Token bucket permits bursts by design, sliding window smooths them, and GCRA gives you sliding-window fairness while storing a single timestamp per key rather than a list.
- Every distributed counter trades accuracy against round trips. Exact counting means synchronous coordination on every request; approximate counting means local decisions with periodic reconciliation and a known overshoot.
- Decide fail-open versus fail-closed per policy, not globally. Quota enforcement should usually fail open; abuse prevention and authentication should not.
The algorithms, and what each one is actually for
Four algorithms cover essentially everything in production. They differ in what they store per key, how they treat bursts, and how much memory and coordination they need.
Fixed window. Count requests in the current clock interval, reset at the boundary. One integer per key, one increment per request, and it is the cheapest thing that could possibly work.
Its flaw is well known and worth stating precisely: a client can send a full window’s allowance at the end of one window and another full allowance at the start of the next, producing twice the intended rate across a short span straddling the boundary. For quota enforcement where the window is a billing month, that flaw is irrelevant. For backend protection where the window is a second, it is the whole problem.
Token bucket. A bucket holds up to N tokens and refills at a constant rate. Each request takes a token; if the bucket is empty the request is rejected. State is a token count and a last-refill timestamp, and refill is computed lazily on read rather than by a background process.
This is the right default for backend protection, because it separates two things that fixed window conflates: the sustained rate, which is the refill rate, and the burst allowance, which is the bucket size. A client that has been quiet accumulates tokens and can spend them in a burst, which is usually what you want, because a quiet client suddenly sending twenty requests is normal behaviour and not abuse. If bursts are the thing you are trying to prevent, set the bucket size equal to one interval’s refill and you have removed the burst allowance deliberately.
Leaky bucket is the same shape with the opposite emphasis: requests enter a queue that drains at a constant rate, so output is perfectly smooth and bursts are delayed rather than rejected. Useful when the thing downstream genuinely cannot absorb variance, such as a third-party API with a hard rate ceiling you must not exceed. Less useful for user-facing traffic, because queueing a request is often worse for the user than rejecting it.
Sliding window log. Store the timestamp of every request in the window, drop timestamps older than the window, and count what remains. Exact, in the sense that it precisely answers “how many requests in the last sixty seconds” at every instant.
The cost is memory proportional to the limit itself, per key. A limit of ten thousand requests per minute means up to ten thousand timestamps per key, and if your key is a customer ID and you have many customers, that adds up quickly. It also means the work per request grows with the limit, since you are trimming a list. Correct, and it does not scale to high limits or many keys.
Sliding window counter. The practical compromise. Keep counters for the current and previous fixed windows, and estimate the sliding count by weighting the previous window by how much of it still overlaps the sliding period. Two integers per key, constant work per request, and it removes the boundary spike that makes fixed window unusable for short intervals.
It is an approximation, and the direction of error is worth knowing: it assumes the previous window’s requests were spread evenly, so a client whose requests were clustered at the start of the previous window is penalised slightly more than it should be, and one clustered at the end gets slightly more headroom. For nearly every real use, that error is irrelevant and the memory saving is not.
GCRA, the generic cell rate algorithm. Borrowed from ATM networking and underused in HTTP. Rather than counting requests, it stores a single timestamp per key called the theoretical arrival time, which is when the next request would be allowed if traffic were perfectly paced. A request is allowed if now is not too far before that time, where “too far” is your burst tolerance.
The properties are genuinely attractive. It gives you sliding-window smoothness without a log, with exactly one value stored per key regardless of the limit. It has a configurable burst tolerance like token bucket. And because the stored value is a timestamp rather than a count, the whole check-and-update is one compare-and-set operation, which makes it unusually clean to implement atomically in a shared store. It also tells you, for free, exactly how long a rejected client should wait, which is the Retry-After value you should have been sending anyway.
Its downside is that it is harder to explain. “You have one hundred requests per minute” is a sentence a customer understands. “Your requests must be paced at an average interval with a burst tolerance” is not, which is why GCRA is a better fit for internal protection than for a published commercial quota.
| Algorithm | State per key | Burst behaviour | Best used for |
|---|---|---|---|
| Fixed window | One counter | Permits double rate at window boundary | Long-window quotas where the boundary spike does not matter |
| Token bucket | Count plus timestamp | Burst up to bucket size, then sustained refill rate | Backend protection where occasional bursts are legitimate |
| Leaky bucket | Queue or level plus timestamp | Smooths by delaying rather than rejecting | Respecting a hard downstream rate ceiling |
| Sliding window log | One timestamp per request | Exact, no boundary artefacts | Low limits where precision genuinely matters |
| Sliding window counter | Two counters | Smooth, with a small estimation error | General-purpose default for short windows |
| GCRA | One timestamp | Configurable tolerance, perfectly paced otherwise | High key counts, and anywhere you want an accurate Retry-After |
Distributed counters: the accuracy you are actually buying
Once the counter is shared, every design is a point on one axis: how much coordination happens per request.
Synchronous shared counter. Every request performs an atomic operation against one store, typically Redis, and the store is the single source of truth. This is exact, in the sense that no two requests can both observe the same remaining allowance and both be admitted, provided the operation really is atomic. In Redis that means a Lua script or a single command, not a read followed by a write.
You pay one network round trip per request, and you have made that store’s availability your API’s availability. The round trip is the honest cost, and it is why rate limiting is frequently the largest single contributor to a gateway’s added latency.
Asynchronous local counting with periodic sync. Each node counts locally and reports to a shared store on an interval, receiving back an estimate of global usage. No round trip on the request path, and the accuracy is bounded by what can happen between syncs. The overshoot is roughly the number of nodes multiplied by what one node can admit in one sync interval, which is a number you can compute and decide to accept.
This is the right design for high-volume protection where approximate is fine, and the wrong design for a hard commercial quota where a customer will notice being allowed over their limit and then charged for it.
Local decisions with a shared budget lease. A middle path: each node requests an allocation of the global budget from the coordinator, spends it locally, and returns for more. Round trips happen per lease rather than per request. Accuracy is good and the failure mode is graceful, since a node holding an unspent lease keeps working briefly without the coordinator. The complexity is in allocating fairly between nodes with unequal traffic, which is where implementations get subtle.
Edge-native counting. Rate limiting at a content delivery network’s edge, where the counter lives in the same point of presence as the request. Very fast and it rejects abuse before it consumes your origin bandwidth, which is the real value. The catch is that “global” limits across points of presence are eventually consistent by necessity, because strong consistency across regions on every request would defeat the purpose. Read every edge vendor’s consistency description carefully, because the words “global rate limit” hide a lot of variance.
The rule that follows from all of this: pick your accuracy requirement before picking an implementation. Abuse prevention tolerates a great deal of overshoot because the attacker is being rejected either way. Commercial quota enforcement tolerates almost none, because the consequence of error is a billing dispute. Backend protection sits between and usually needs less precision than teams assume.
Needs first-hand data: Configure the same limit under each strategy, drive a fixed workload across a known number of nodes, and measure actual admitted requests against the configured limit, plus added p99 latency. The overshoot under async local counting and the latency under synchronous counting are both predictable in shape and unpublished in magnitude, and one table would settle most architecture arguments in this category.
What happens under node failure and partition
This is the part that decides whether your rate limiter helps or hurts during an incident, and it is the part evaluation processes skip.
The store becomes slow. More common than the store going down, and worse, because a slow dependency in the request path turns into queued requests, exhausted connection pools and cascading timeouts. If your rate limit check has no timeout, you have built a system where a slow Redis takes down a healthy API. Set a timeout on the rate limit check that is a small fraction of your request budget, and decide explicitly what happens when it fires.
The store becomes unreachable. Now the fail-open versus fail-closed decision is being made, and it should have been made deliberately in advance.
Fail open means admitting the request when you cannot check. Correct for commercial quota enforcement, because briefly letting a customer exceed their plan is a billing adjustment, while rejecting every paying customer because a cache is down is an outage. Also correct for most backend protection, on the reasoning that your backend has other defences and total rejection is guaranteed damage against a risk of damage.
Fail closed means rejecting when you cannot check. Correct for abuse prevention and for anything adjacent to authentication, where the whole point is that the unverified case is the dangerous one. If your limiter is what stands between a login endpoint and a credential-stuffing run, failing open during an outage is an invitation.
The important thing is that these should be configured per policy. A system with one global setting forces you to be wrong about half your policies.
A partition splits the cluster. Each side counts independently, so the effective global limit becomes the sum of what each partition admits. If you are using a shared store with a single primary, one side loses access entirely and its behaviour is governed by your fail-open decision. If you are using a replicated or multi-region store, both sides keep working and both count only their own traffic. Either way, during a partition you are not enforcing a global limit, and no configuration fixes that. What you can do is know it, and make sure the consequence is a temporary overshoot rather than a hard failure.
A node restarts. Local counters are gone. With a shared store this is a non-event. With async local counting, the node starts from zero and can admit a burst before its first sync. With lease-based allocation, the node’s unspent lease is lost, which is safe but wasteful. Worth checking in any implementation where deploys are frequent, because a rolling deploy is a sequence of counter resets.
Redis

Redis is the default shared counter for rate limiting and, for most teams, the correct answer. Its value here is not speed in the abstract but that Lua scripts execute atomically, which makes a correct check-and-update a single round trip rather than a read-modify-write race. Every rate limiting algorithm above has a well-known Redis implementation, and there are mature libraries for GCRA specifically.
Pros
- Atomic Lua script execution makes correct distributed rate limiting genuinely straightforward to implement
- Every algorithm here maps cleanly onto Redis data structures, so you are not constrained to whatever a vendor implemented
- You control the failure behaviour completely, including per-policy fail-open and fail-closed decisions and check timeouts
- Almost certainly already in your stack, so no new vendor relationship or procurement
Cons
- One network round trip per request on the hot path, which is often the largest single component of added gateway latency
- You now own availability of something in the request path of every API call, including its failover behaviour
- Cluster mode complicates atomic multi-key operations, so keys must be designed to hash to one slot or the scripts break in subtle ways
Best for: Teams that already operate Redis reliably and want exact control over the algorithm and the failure behaviour.
Pricing: Open source with no licence cost when self-hosted; managed offerings meter on memory, throughput or connection count depending on the provider.
Upstash

Upstash is Redis exposed over HTTP with serverless pricing, which solves a specific and real problem: edge and serverless runtimes often cannot hold a persistent TCP connection to a conventional Redis, and connection-per-invocation against a normal Redis exhausts it. It ships a rate limiting library implementing the common algorithms, so the integration is a few lines rather than a Lua script.
Pros
- HTTP interface works from edge runtimes and serverless functions where a persistent Redis connection is not available
- Request-based pricing fits spiky and low-volume workloads where a provisioned instance is mostly idle
- The rate limiting library implements the standard algorithms directly, including sliding window variants, with sensible defaults
- Global replication options put the counter closer to the request without you operating regions
Cons
- HTTP adds protocol overhead versus the Redis wire protocol, so per-check latency is worse than a co-located Redis
- Per-request pricing means the rate limiter has a marginal cost on every API call, which inverts badly at high volume
- Globally replicated counters are eventually consistent, so a strict global limit is not what you are getting
Best for: Serverless and edge applications that need a shared counter and cannot hold a conventional Redis connection.
Pricing: Metered per request with storage and bandwidth components, plus provisioned options for predictable high-volume workloads.
Cloudflare
Cloudflare’s rate limiting runs at the edge, which changes what the tool is for. Rejecting abusive traffic in the point of presence nearest the client means it never touches your origin, never consumes your bandwidth and never occupies a connection on your gateway. That is a categorically different benefit from protecting a backend behind your own proxy, and for abuse and volumetric protection it is the right layer.
Pros
- Rejection happens before traffic reaches your infrastructure, which is the only layer that protects against volumetric abuse
- Rules can match on rich request attributes including path, method, headers and bot signals, not just client address
- No infrastructure of your own in the request path, and no counter for you to operate
- Integrates with the rest of the edge security stack, so rate limiting is one control among several rather than a standalone system
Cons
- Cross-point-of-presence counting is eventually consistent by design, so a precise global limit is not achievable at this layer
- Rules live in vendor configuration rather than your application repository, which splits your policy across two systems
- Not suited to per-customer commercial quotas, because the edge does not have your entitlement data
Best for: Public-facing APIs that need volumetric and abuse protection before traffic reaches origin, alongside a separate quota mechanism further in.
Pricing: Included in plan tiers with rule count and evaluation volume as the metered dimensions, with higher tiers unlocking more expressive matching.
Kong Gateway

If you already run a gateway, rate limiting there is usually the right place for it, and Kong’s plugin is the most widely deployed implementation in this list. It supports several storage strategies, including node-local, Redis-backed and a cluster-coordinated mode, which is precisely the accuracy-versus-latency axis described above expressed as a configuration option.
Pros
- Enforcement sits where consumer identity is already resolved, so per-consumer limits need no additional plumbing
- Multiple storage strategies let you trade accuracy against latency per route rather than globally
- Fail-open behaviour when the store is unreachable is configurable rather than assumed
- Limits are declared in the same configuration as routes and auth, so policy lives in one reviewable place
Cons
- The open source implementation has weaker distributed accuracy than the commercial advanced plugin, and the gap is easy to miss until it matters
- Redis-backed mode puts Redis in the request path of every rate-limited route, which is a dependency teams add without planning for it
- Only protects traffic that traverses the gateway, so internal service-to-service calls need separate handling
Best for: Teams already running Kong who want per-consumer limits enforced where identity is already known.
Pricing: Included in the open source gateway at no licence cost; the more accurate advanced implementation is part of the commercial tier.
Apache APISIX

APISIX ships several rate limiting plugins covering request count, concurrency and connection limits, all available without a licence, which is the main reason to choose it over the open build of a commercial gateway. Storage is local or Redis-backed per plugin instance, and because the plugin set is open the implementation is readable, which matters more than usual when you are reasoning about accuracy.
Pros
- Multiple limiting plugins covering request rate, concurrency and connections, all in the open build with no licence cost
- Redis and Redis cluster backing available without a commercial tier, so accurate distributed limiting is not gated
- Plugin source is readable, so you can verify the algorithm and the failure behaviour rather than trusting documentation
- Limits can attach at route, service or consumer scope, which covers both protection and quota use cases
Cons
- Configuration spread across several plugins with overlapping purposes, and choosing the right one is less obvious than it should be
- etcd is required for the gateway itself, so you are operating two stateful systems if you also add Redis for limits
- Documentation on the precise consistency guarantees of each storage mode is thinner than the implementation deserves
Best for: Teams running APISIX who want accurate distributed limiting without paying for a commercial tier.
Pricing: Apache-licensed with no licence cost; the operational cost is etcd plus whatever store backs distributed counting.
Unkey
Unkey approaches this from the API key side rather than the proxy side: it manages keys, and rate limits and quotas are properties of a key rather than of a route. That inversion suits products where limits are commercial entitlements attached to a customer, and it means verification and limiting are one call rather than two systems that must agree about identity.
Pros
- Limits are attached to keys rather than routes, which matches how commercial quotas are actually defined
- Key verification and limit check happen in one operation, removing a class of bugs where auth and limiting disagree about who the caller is
- Designed for edge and serverless runtimes, with a latency model built around distributed verification rather than a central store
- Removes the need to build key issuance, rotation and revocation yourself, which is usually the larger part of this problem
Cons
- Only covers traffic authenticated with its keys, so it is not a general-purpose limiter for unauthenticated or internal traffic
- A third-party dependency in the authentication path, which is a meaningful availability decision
- Younger product with a smaller operational track record than the incumbents here
Best for: Product teams shipping a public API where rate limits are per-customer entitlements and key management is not yet built.
Pricing: Metered on key verifications with tiered plans, plus limits on key count and analytics retention by tier.
How to choose
First, separate your three problems. Write down which limits exist for backend protection, which for commercial quota, and which for abuse prevention. They will want different algorithms, different accuracy and different failure behaviour, and a single mechanism cannot serve all three well.
Second, pick the enforcement layer per problem. Abuse protection belongs at the edge, because rejecting at origin still costs you the bandwidth and the connection. Commercial quota belongs where consumer identity is resolved, which is your gateway or your key management service. Backend protection belongs closest to the thing being protected, which is sometimes the service itself.
Third, choose the algorithm from the requirement, not the default. Token bucket for protection where bursts are legitimate. Sliding window counter for short windows where boundary spikes matter. GCRA when you have very many keys or want an accurate Retry-After. Fixed window for monthly quotas, where it is entirely adequate and cheapest.
Fourth, decide fail-open and fail-closed per policy, and write it down. Then test it, by actually taking the store away in a staging environment under load and watching what happens. This is the single most valuable hour in the entire evaluation and almost nobody spends it.
Fifth, return the right response. A 429 with a Retry-After header and headers describing the limit, the remaining allowance and the reset time. Clients that know when to retry produce far less load than clients that guess, and a rate limiter that causes a retry storm has made things worse. If you expose limits to customers, the values you return are part of your API contract and changing them belongs in your versioning and deprecation process.
Whatever you pick, instrument it. You need to see rejection rate by consumer and by policy, or you will never know whether a limit is protecting you or quietly breaking a customer’s integration. That data belongs alongside the rest of your API analytics, and the symptom of a limit set too tight usually shows up first in API monitoring as an error rate you cannot explain.
Needs first-hand data: Take one production-shaped workload and run it against a Redis-backed limiter while injecting store latency in steps, then while making the store unreachable entirely. Record gateway p99, error rate and admitted requests at each step under both fail-open and fail-closed configuration. Everyone reasons about this behaviour and nobody measures it.
Frequently asked questions
Should rate limiting live in the gateway or the application?
The gateway, for anything expressed per consumer, because that is where the consumer identity is already resolved and where rejection is cheapest. Applications should still protect their own expensive operations, particularly ones that are costly for reasons the gateway cannot see, such as a query whose cost depends on its parameters. Both layers is normal; only the application layer means unauthenticated abuse reaches your services.
What is the difference between rate limiting and a quota?
Timescale and intent. A rate limit is about instantaneous pressure, measured in seconds, and exists to protect something. A quota is about consumption over a billing period, measured in days or months, and exists to enforce a commercial agreement. They need different algorithms and different failure behaviour, and conflating them is why so many implementations are simultaneously too strict and too loose.
Which algorithm should I use if I only implement one?
Sliding window counter, backed by a shared store. It has no boundary spike, needs constant memory per key, is straightforward to implement correctly, and is easy enough to explain to a customer. Token bucket is a better choice if legitimate clients burst, and GCRA is better if you have a very large number of keys, but the sliding window counter is the safest single choice.
Do I need Redis for distributed rate limiting?
You need something shared, and Redis is the usual answer because atomic scripting makes correctness easy. The alternatives are a dedicated rate limit service with its own store, an edge provider’s counters, or async local counting with periodic reconciliation if you can accept bounded overshoot. What does not work is per-replica counters with no coordination, which is what you have by default and which stops being a limit the moment you scale out.
How should I rate limit by client address behind a proxy?
Carefully, and usually not as your primary key. The address you see is the proxy’s unless you read a forwarded header, and forwarded headers are client-controlled unless your edge overwrites them, which makes naive implementations trivially bypassable. Where you do use addresses, remember that a single corporate or mobile carrier address represents many users, and that address-based limits on IPv6 need a prefix rather than a full address. Authenticated identity is always the better key when you have one.
What should a rate limited response actually contain?
Status 429, a Retry-After header with a concrete value rather than a guess, and headers describing the limit, the remaining allowance and the reset time so clients can pace themselves before being rejected. Include enough in the body to tell a developer which policy they hit, because “too many requests” with no policy name is unactionable when several limits apply to the same endpoint.
Related reading
- Best API management platforms — where rate limiting sits in the wider gateway and platform decision.
- Best API gateways — the shared-state dependency every gateway introduces, and how each handles it.
- Best API gateways for Kubernetes — why per-replica limits are the default failure in a cluster.
- Best open source API gateways — which projects gate accurate distributed limiting behind a commercial tier.
- Best managed Redis providers — running the store your limiter depends on without owning its failover.
- Best API analytics tools — seeing rejection rates per consumer, which is how you learn a limit is wrong.